Model comparison

o1 vs o4-mini

o1 and o4-mini score almost the same on the Noometry Index (40.9 vs 41.6), so choose on price, context window or the category you care about most.

Last verified . 41 shared benchmarks.

o1 OpenAI

40.9

Rank #143 Confirmed

o4-mini OpenAI

41.6

Rank #132 Confirmed

Summary

  • They share 41 benchmarks with published results for both. o1 scores higher in 5 categories and o4-mini in 5 categories; 9 gaps are clear of the uncertainty.
  • The widest gap is in agentic & tool use, where o4-mini leads 32.6 to 24.6.
  • The biggest single-benchmark swing is ARC-AGI-1: 30.7% for o1 and 58.7% for o4-mini.
  • o4-mini is cheaper at $1.10 / $4.40 per million input/output tokens, against $15 / $60 for o1.

Side by side

o1 and o4-mini specifications
o1o4-mini
ProviderOpenAIOpenAI
Noometry Index40.941.6
Released2024-09-122025-04-16
WeightsProprietaryProprietary
Context window200K200K
Max output100K100K
Input $ / M tokens$15$1.10
Output $ / M tokens$60$4.40
Results tracked5260

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding o1 leads

o1: 46.1 (#70), o4-mini: 40.9 (#127)

Coding benchmarks
Benchmarko1o4-mini
Aider Polyglot61.7%72%
WeirdML47.6%52.6%
LMArena Coding13671368
CadEval56%62%
SWE-bench Verified (bash only)—45%
GSO—3.6%
LiveBench Coding69.7%—
ALE-Bench—826.17
AlgoTune—1.72
HumanEval+89%—
MBPP+80.2%—

Agentic & Tool Use o4-mini leads

o1: 24.6 (#117), o4-mini: 32.6 (#61)

Agentic & Tool Use benchmarks
Benchmarko1o4-mini
METR Time Horizons51.1%63.9%
Berkeley Function Calling Leaderboard—53.2%
GDPval—25.3%
Cybench10%—

Reasoning o1 leads

o1: 27.9 (#111), o4-mini: 24.6 (#162)

Reasoning benchmarks
Benchmarko1o4-mini
SimpleBench41.7%38.7%
ARC-AGI-130.7%58.7%
Chess Puzzles15%26%
EnigmaEval5.7%9.2%
LMArena Hard Prompts13711351
DTBench74.7%77.6%
LMCA22.3%26.5%
Epoch Capabilities Index141.91145.64
ARC-AGI-2—6.1%
Kagi LLM Benchmark—67.6%
CritPt—0.6%
LiveBench Reasoning91.6%—
Mystery Game Puzzles—5%
LiveBench Data Analysis65.5%—
ForecastBench—61.8
LiveBench75.7%—

Math o4-mini leads

o1: 36.1 (#175), o4-mini: 40.8 (#89)

Knowledge o4-mini leads

o1: 41.5 (#110), o4-mini: 43.6 (#91)

Knowledge benchmarks
Benchmarko1o4-mini
GPQA Diamond76.8%79.6%
Humanity's Last Exam8%18.1%
SimpleQA Verified41.1%19.6%
Confabulations11.7%15.8%
LMArena Expert13611343
MMLU-Pro—82%
Vectara Hallucination Rate—18.6%
GPQA (HELM)—73.5%

Multimodal o4-mini leads

o1: 34.2 (#93), o4-mini: 40.2 (#49)

Multimodal benchmarks
Benchmarko1o4-mini
LMArena Vision11681194
GeoBench80%64%
VPCT37%57.5%
SpatialViz-Bench41.4%—

Multilingual o1 leads

o1: 48.6 (#142), o4-mini: 47.0 (#154)

Multilingual benchmarks
Benchmarko1o4-mini
LMArena Non-English13581337
LMArena Chinese13941354
LMArena French13441364
LMArena German13371336
LMArena Japanese13461308
LMArena Korean13961312
LMArena Russian13561334
LMArena Spanish13451347

Instruction Following Too close to call

o1: 74.8 (#86), o4-mini: 75.2 (#68)

Instruction Following benchmarks
Benchmarko1o4-mini
LMArena Instruction Following13671321
LiveBench Instruction Following81.5%—
IFEval—92.8%

Long Context o1 leads

o1: 50.3 (#9), o4-mini: 45.5 (#33)

Long Context benchmarks
Benchmarko1o4-mini
Fiction.LiveBench83.3%77.8%
LMArena Longer Query13781315

Writing & Preference o1 leads

o1: 55.6 (#144), o4-mini: 54.0 (#152)

Writing & Preference benchmarks
Benchmarko1o4-mini
LMArena Text13661353
LMArena Creative Writing13481294
Short-Story Creative Writing70.2%75%
LMArena Multi-Turn13691350
WildBench—85.4%
LiveBench Language65.4%—

Frequently asked questions

Is o1 better than o4-mini?

o1 and o4-mini score almost the same on the Noometry Index (40.9 vs 41.6), so choose on price, context window or the category you care about most.

Which is cheaper, o1 or o4-mini?

o4-mini is cheaper. It lists at $1.10 per million input tokens and $4.40 per million output tokens; o1 lists at $15 and $60.

Is o1 or o4-mini better for coding?

o1 scores higher on coding benchmarks: 46.1 versus 40.9 in the Noometry coding category.

Which has the bigger context window?

Both accept 200K tokens.

How many benchmarks do o1 and o4-mini share?

41 benchmarks have published results for both models. o1 has 52 scored results on Noometry and o4-mini has 60.

Related comparisons

Go deeper