Model comparison

o1 vs o3

o3 is the stronger model overall, scoring 47.5 to 40.9 on the Noometry Index.

Last verified . 41 shared benchmarks.

o1 OpenAI

40.9

Rank #143 Confirmed

o3 OpenAI

47.5

Rank #61 Confirmed

Summary

  • They share 41 benchmarks with published results for both. o1 scores higher in 1 category and o3 in 9 categories; 9 gaps are clear of the uncertainty.
  • The widest gap is in math, where o3 leads 50.2 to 36.1.
  • The biggest single-benchmark swing is ARC-AGI-1: 30.7% for o1 and 60.8% for o3.
  • o3 is cheaper at $2 / $8 per million input/output tokens, against $15 / $60 for o1.

Side by side

o1 and o3 specifications
o1o3
ProviderOpenAIOpenAI
Noometry Index40.947.5
Released2024-09-122025-04-16
WeightsProprietaryProprietary
Context window200K200K
Max output100K100K
Input $ / M tokens$15$2
Output $ / M tokens$60$8
Results tracked5263

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

o1: 46.1 (#70), o3: 46.8 (#64)

Coding benchmarks
Benchmarko1o3
Aider Polyglot61.7%81.3%
WeirdML47.6%52.4%
LMArena Coding13671408
CadEval56%74%
SWE-bench Verified—62.3%
SWE-bench Verified (bash only)—58.4%
GSO—8.8%
LiveBench Coding69.7%—
ALE-Bench—933.55
HumanEval+89%—
MBPP+80.2%—

Agentic & Tool Use o3 leads

o1: 24.6 (#117), o3: 34.5 (#44)

Agentic & Tool Use benchmarks
Benchmarko1o3
METR Time Horizons51.1%65.4%
Berkeley Function Calling Leaderboard—63%
GDPval—30.8%
Cybench10%—
DeepResearch Bench—45.2%
OSWorld—23%
LMArena Search—1144

Reasoning o3 leads

o1: 27.9 (#111), o3: 32.0 (#78)

Reasoning benchmarks
Benchmarko1o3
SimpleBench41.7%53.1%
ARC-AGI-130.7%60.8%
Chess Puzzles15%38%
EnigmaEval5.7%13.1%
LMArena Hard Prompts13711402
DTBench74.7%84.8%
LMCA22.3%39.7%
Epoch Capabilities Index141.91146.86
ARC-AGI-2—6.5%
Kagi LLM Benchmark—67.6%
CritPt—1.4%
LiveBench Reasoning91.6%—
Mystery Game Puzzles—29%
LiveBench Data Analysis65.5%—
ForecastBench—62.5
LiveBench75.7%—

Math o3 leads

o1: 36.1 (#175), o3: 50.2 (#58)

Math benchmarks
Benchmarko1o3
FrontierMath (Tiers 1-3)14.7%33.3%
OTIS Mock AIME 2024-202573.3%84.4%
LMArena Math13881426
MATH Level 594.7%97.8%
FrontierMath (Feb 2025 set)9.3%18.7%
Omni-MATH—71.4%
LiveBench Math80.3%—
FrontierMath Tier 4 (v1)—2.1%

Knowledge o3 leads

o1: 41.5 (#110), o3: 54.6 (#52)

Knowledge benchmarks
Benchmarko1o3
GPQA Diamond76.8%81.8%
Humanity's Last Exam8%20.3%
SimpleQA Verified41.1%49.4%
Confabulations11.7%14.4%
LMArena Expert13611402
MMLU-Pro—85.9%
GPQA (HELM)—75.3%

Multimodal o3 leads

o1: 34.2 (#93), o3: 41.4 (#36)

Multimodal benchmarks
Benchmarko1o3
LMArena Vision11681214
GeoBench80%74%
VPCT37%52%
SpatialViz-Bench41.4%—

Multilingual o3 leads

o1: 48.6 (#142), o3: 51.7 (#105)

Multilingual benchmarks
Benchmarko1o3
LMArena Non-English13581401
LMArena Chinese13941437
LMArena French13441430
LMArena German13371420
LMArena Japanese13461403
LMArena Korean13961370
LMArena Russian13561406
LMArena Spanish13451395

Instruction Following o1 leads

o1: 74.8 (#86), o3: 72.8 (#127)

Instruction Following benchmarks
Benchmarko1o3
LMArena Instruction Following13671368
LiveBench Instruction Following81.5%—
IFEval—86.9%

Long Context o3 leads

o1: 50.3 (#9), o3: 53.3 (#6)

Long Context benchmarks
Benchmarko1o3
Fiction.LiveBench83.3%88.9%
LMArena Longer Query13781372
CL-bench—17.8%

Writing & Preference o3 leads

o1: 55.6 (#144), o3: 63.5 (#64)

Writing & Preference benchmarks
Benchmarko1o3
LMArena Text13661410
LMArena Creative Writing13481359
Short-Story Creative Writing70.2%83.9%
LMArena Multi-Turn13691405
EQ-Bench Creative Writing—1676
WildBench—86.1%
LiveBench Language65.4%—

Frequently asked questions

Is o1 better than o3?

o3 is the stronger model overall, scoring 47.5 to 40.9 on the Noometry Index.

Which is cheaper, o1 or o3?

o3 is cheaper. It lists at $2 per million input tokens and $8 per million output tokens; o1 lists at $15 and $60.

Is o1 or o3 better for coding?

They score almost the same on coding (46.1 vs 46.8); test both on your own repository before choosing.

Which has the bigger context window?

Both accept 200K tokens.

How many benchmarks do o1 and o3 share?

41 benchmarks have published results for both models. o1 has 52 scored results on Noometry and o3 has 63.

Related comparisons

Go deeper