Model comparison

GPT-4.1 vs o1-mini

GPT-4.1 is the stronger model overall, scoring 35.9 to 34.0 on the Noometry Index.

Last verified . 27 shared benchmarks.

GPT-4.1 OpenAI

35.9

Rank #219 Confirmed

o1-mini OpenAI

34.0

Rank #235 Confirmed

Summary

  • They share 27 benchmarks with published results for both. GPT-4.1 scores higher in 6 categories and o1-mini in 3 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in math, where o1-mini leads 35.4 to 22.3.
  • The biggest single-benchmark swing is Aider Polyglot: 52.4% for GPT-4.1 and 32.9% for o1-mini.

Side by side

GPT-4.1 and o1-mini specifications
GPT-4.1o1-mini
ProviderOpenAIOpenAI
Noometry Index35.934.0
Released2025-04-142024-09-12
WeightsProprietaryProprietary
Context window1.05M—
Max output33K—
Input $ / M tokens$2—
Output $ / M tokens$8—
Results tracked5239

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding o1-mini leads

GPT-4.1: 34.4 (#238), o1-mini: 35.5 (#224)

Coding benchmarks
BenchmarkGPT-4.1o1-mini
Aider Polyglot52.4%32.9%
WeirdML39%36.3%
LMArena Coding13911362
SWE-bench Verified48.5%—
SWE-bench Verified (bash only)39.6%—
LiveBench Coding—48%
CadEval42%—
ALE-Bench558.1—
HumanEval+—89%
MBPP+—78.8%

Agentic & Tool Use GPT-4.1 leads

GPT-4.1: 34.7 (#43), o1-mini: 24.6 (#118)

Agentic & Tool Use benchmarks
BenchmarkGPT-4.1o1-mini
Berkeley Function Calling Leaderboard54%—
Cybench—10%

Reasoning GPT-4.1 leads

GPT-4.1: 11.7 (#339), o1-mini: 8.8 (#346)

Reasoning benchmarks
BenchmarkGPT-4.1o1-mini
ARC-AGI-20.4%0.8%
SimpleBench27%18.1%
ARC-AGI-15.5%14%
LMArena Hard Prompts13841333
Epoch Capabilities Index136.78135.82
Kagi LLM Benchmark52.3%—
Chess Puzzles6%—
EnigmaEval2.2%—
LiveBench Reasoning—72.3%
DTBench68.3%—
LiveBench Data Analysis—57.9%
LMCA25.6%—
ForecastBench61.5—
LiveBench—57.8%

Math o1-mini leads

GPT-4.1: 22.3 (#280), o1-mini: 35.4 (#186)

Knowledge GPT-4.1 leads

GPT-4.1: 37.1 (#160), o1-mini: 34.9 (#192)

Knowledge benchmarks
BenchmarkGPT-4.1o1-mini
GPQA Diamond66.9%62.4%
LMArena Expert13641316
Humanity's Last Exam5.4%—
SimpleQA Verified31.1%—
MMLU-Pro81.1%—
Confabulations—18.6%
Vectara Hallucination Rate5.6%—
GPQA (HELM)65.9%—

Multimodal Not comparable

GPT-4.1: 38.2 (#67), o1-mini: —

Multimodal benchmarks
BenchmarkGPT-4.1o1-mini
LMArena Vision1211—
GeoBench72%—

Multilingual GPT-4.1 leads

GPT-4.1: 49.4 (#133), o1-mini: 43.6 (#182)

Multilingual benchmarks
BenchmarkGPT-4.1o1-mini
LMArena Non-English13701289
LMArena Chinese13821314
LMArena French13821293
LMArena German13811278
LMArena Japanese13191245
LMArena Korean13391223
LMArena Russian13771283
LMArena Spanish13761303

Instruction Following GPT-4.1 leads

GPT-4.1: 71.3 (#153), o1-mini: 66.7 (#206)

Instruction Following benchmarks
BenchmarkGPT-4.1o1-mini
LMArena Instruction Following13671304
LiveBench Instruction Following—65.4%
IFEval83.8%—

Long Context Too close to call

GPT-4.1: 40.0 (#163), o1-mini: 40.1 (#161)

Long Context benchmarks
BenchmarkGPT-4.1o1-mini
LMArena Longer Query13851320
Fiction.LiveBench63.9%—

Writing & Preference GPT-4.1 leads

GPT-4.1: 57.6 (#125), o1-mini: 48.4 (#202)

Writing & Preference benchmarks
BenchmarkGPT-4.1o1-mini
LMArena Text13831317
LMArena Creative Writing13631244
LMArena Multi-Turn13981314
Short-Story Creative Writing—64.9%
EQ-Bench Creative Writing1420—
WildBench85.4%—
LiveBench Language—40.9%

Frequently asked questions

Is GPT-4.1 better than o1-mini?

GPT-4.1 is the stronger model overall, scoring 35.9 to 34.0 on the Noometry Index.

Is GPT-4.1 or o1-mini better for coding?

o1-mini scores higher on coding benchmarks: 35.5 versus 34.4 in the Noometry coding category.

How many benchmarks do GPT-4.1 and o1-mini share?

27 benchmarks have published results for both models. GPT-4.1 has 52 scored results on Noometry and o1-mini has 39.

Related comparisons

Go deeper