Model comparison

o3 vs Step 5 Preview

o3 and Step 5 Preview score almost the same on the Noometry Index (47.5 vs 47.9), so choose on price, context window or the category you care about most.

Last verified . 15 shared benchmarks.

o3 OpenAI

47.5

Rank #61 Confirmed

Step 5 Preview StepFun

47.9

Rank #58 Confirmed

Summary

  • They share 15 benchmarks with published results for both. o3 scores higher in 5 categories and Step 5 Preview in 4 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where o3 leads 54.6 to 41.2.
  • The biggest single-benchmark swing is CritPt: 1.4% for o3 and 20.9% for Step 5 Preview.
  • Step 5 Preview is cheaper at $1 / $2.70 per million input/output tokens, against $2 / $8 for o3.
  • Step 5 Preview accepts more context: 1.02M tokens versus 200K.

Side by side

o3 and Step 5 Preview specifications
o3Step 5 Preview
ProviderOpenAIStepFun
Noometry Index47.547.9
Released2025-04-162026-09-16
WeightsProprietaryProprietary
Context window200K1.02M
Max output100K66K
Input $ / M tokens$2$1
Output $ / M tokens$8$2.70
Results tracked6318

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Step 5 Preview leads

o3: 46.8 (#64), Step 5 Preview: 51.8 (#36)

Coding benchmarks
Benchmarko3Step 5 Preview
LMArena Coding14081480
SWE-bench Verified62.3%—
SWE-bench Verified (bash only)58.4%—
Aider Polyglot81.3%—
LMArena WebDev—1564
SciCode—58.9%
GSO8.8%—
WeirdML52.4%—
CadEval74%—
ALE-Bench933.55—

Agentic & Tool Use Not comparable

o3: 34.5 (#44), Step 5 Preview: —

Agentic & Tool Use benchmarks
Benchmarko3Step 5 Preview
Berkeley Function Calling Leaderboard63%—
GDPval30.8%—
DeepResearch Bench45.2%—
OSWorld23%—
LMArena Search1144—
METR Time Horizons65.4%—

Reasoning Step 5 Preview leads

o3: 32.0 (#78), Step 5 Preview: 40.0 (#57)

Reasoning benchmarks
Benchmarko3Step 5 Preview
CritPt1.4%20.9%
LMArena Hard Prompts14021465
ARC-AGI-26.5%—
SimpleBench53.1%—
Kagi LLM Benchmark67.6%—
ARC-AGI-160.8%—
Chess Puzzles38%—
EnigmaEval13.1%—
Mystery Game Puzzles29%—
DTBench84.8%—
LMCA39.7%—
Epoch Capabilities Index146.86—
ForecastBench62.5—

Math o3 leads

o3: 50.2 (#58), Step 5 Preview: 46.1 (#72)

Math benchmarks
Benchmarko3Step 5 Preview
LMArena Math14261470
FrontierMath (Tiers 1-3)33.3%—
OTIS Mock AIME 2024-202584.4%—
ProofBench—42%
Omni-MATH71.4%—
MATH Level 597.8%—
FrontierMath (Feb 2025 set)18.7%—
FrontierMath Tier 4 (v1)2.1%—

Knowledge o3 leads

o3: 54.6 (#52), Step 5 Preview: 41.2 (#112)

Knowledge benchmarks
Benchmarko3Step 5 Preview
LMArena Expert14021470
GPQA Diamond81.8%—
Humanity's Last Exam20.3%—
SimpleQA Verified49.4%—
MMLU-Pro85.9%—
Confabulations14.4%—
GPQA (HELM)75.3%—

Multimodal Too close to call

o3: 41.4 (#36), Step 5 Preview: 41.0 (#41)

Multimodal benchmarks
Benchmarko3Step 5 Preview
LMArena Vision12141267
GeoBench74%—
VPCT52%—

Multilingual Step 5 Preview leads

o3: 51.7 (#105), Step 5 Preview: 53.8 (#54)

Multilingual benchmarks
Benchmarko3Step 5 Preview
LMArena Non-English14011432
LMArena Chinese14371519
LMArena Russian14061436
LMArena Spanish13951439
LMArena French1430—
LMArena German1420—
LMArena Japanese1403—
LMArena Korean1370—

Instruction Following Step 5 Preview leads

o3: 72.8 (#127), Step 5 Preview: 76.0 (#50)

Instruction Following benchmarks
Benchmarko3Step 5 Preview
LMArena Instruction Following13681444
IFEval86.9%—

Long Context o3 leads

o3: 53.3 (#6), Step 5 Preview: 44.5 (#64)

Long Context benchmarks
Benchmarko3Step 5 Preview
LMArena Longer Query13721455
Fiction.LiveBench88.9%—
CL-bench17.8%—

Writing & Preference Too close to call

o3: 63.5 (#64), Step 5 Preview: 62.8 (#71)

Writing & Preference benchmarks
Benchmarko3Step 5 Preview
LMArena Text14101442
LMArena Creative Writing13591410
LMArena Multi-Turn14051447
Short-Story Creative Writing83.9%—
EQ-Bench Creative Writing1676—
WildBench86.1%—

Frequently asked questions

Is o3 better than Step 5 Preview?

o3 and Step 5 Preview score almost the same on the Noometry Index (47.5 vs 47.9), so choose on price, context window or the category you care about most.

Which is cheaper, o3 or Step 5 Preview?

Step 5 Preview is cheaper. It lists at $1 per million input tokens and $2.70 per million output tokens; o3 lists at $2 and $8.

Is o3 or Step 5 Preview better for coding?

Step 5 Preview scores higher on coding benchmarks: 51.8 versus 46.8 in the Noometry coding category.

Which has the bigger context window?

Step 5 Preview does, with 1.02M tokens against 200K.

How many benchmarks do o3 and Step 5 Preview share?

15 benchmarks have published results for both models. o3 has 63 scored results on Noometry and Step 5 Preview has 18.

Related comparisons

Go deeper