Model comparison

GPT-5.1 vs o3

GPT-5.1 is the stronger model overall, scoring 49.0 to 47.5 on the Noometry Index.

Last verified . 49 shared benchmarks.

GPT-5.1 OpenAI

49.0

Rank #53 Confirmed

o3 OpenAI

47.5

Rank #61 Confirmed

Summary

  • They share 49 benchmarks with published results for both. GPT-5.1 scores higher in 6 categories and o3 in 4 categories; 9 gaps are clear of the uncertainty.
  • The widest gap is in instruction following, where GPT-5.1 leads 83.9 to 72.8.
  • The biggest single-benchmark swing is GPQA (HELM): 44.2% for GPT-5.1 and 75.3% for o3.
  • Both cost about the same: $1.25 input and $10 output per million tokens.
  • GPT-5.1 accepts more context: 400K tokens versus 200K.

Side by side

GPT-5.1 and o3 specifications
GPT-5.1o3
ProviderOpenAIOpenAI
Noometry Index49.047.5
Released2025-11-132025-04-16
WeightsProprietaryProprietary
Context window400K200K
Max output128K100K
Input $ / M tokens$1.25$2
Output $ / M tokens$10$8
Results tracked6363

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

GPT-5.1: 46.4 (#66), o3: 46.8 (#64)

Coding benchmarks
BenchmarkGPT-5.1o3
SWE-bench Verified68%62.3%
SWE-bench Verified (bash only)66%58.4%
GSO13.7%8.8%
WeirdML60.8%52.4%
LMArena Coding14541408
ALE-Bench1,192933.55
Aider Polyglot—81.3%
LMArena WebDev1395—
SciCode43.3%—
LiveBench Coding72.5%—
CadEval—74%

Agentic & Tool Use o3 leads

GPT-5.1: 32.7 (#60), o3: 34.5 (#44)

Agentic & Tool Use benchmarks
BenchmarkGPT-5.1o3
DeepResearch Bench42.8%45.2%
LMArena Search11991144
Terminal-Bench47.6%—
Berkeley Function Calling Leaderboard—63%
GDPval—30.8%
OSWorld—23%
METR Time Horizons—65.4%
Vending-Bench 21,473—

Reasoning GPT-5.1 leads

GPT-5.1: 39.8 (#58), o3: 32.0 (#78)

Reasoning benchmarks
BenchmarkGPT-5.1o3
ARC-AGI-217.6%6.5%
SimpleBench53.2%53.1%
ARC-AGI-172.8%60.8%
CritPt4.9%1.4%
Chess Puzzles32%38%
EnigmaEval11.2%13.1%
LMArena Hard Prompts14571402
Mystery Game Puzzles19%29%
DTBench90.1%84.8%
LMCA43.9%39.7%
Epoch Capabilities Index149.64146.86
ForecastBench58.162.5
Kagi LLM Benchmark—67.6%
LiveBench Reasoning95.8%—
LiveBench Data Analysis72.1%—
LiveBench78.8%—

Math GPT-5.1 leads

GPT-5.1: 52.2 (#51), o3: 50.2 (#58)

Math benchmarks
BenchmarkGPT-5.1o3
OTIS Mock AIME 2024-202588.6%84.4%
Omni-MATH46.4%71.4%
LMArena Math14471426
FrontierMath (Feb 2025 set)31%18.7%
FrontierMath Tier 4 (v1)12.5%2.1%
FrontierMath (Tiers 1-3)—33.3%
LiveBench Math94.5%—
MATH Level 5—97.8%

Knowledge o3 leads

GPT-5.1: 50.6 (#71), o3: 54.6 (#52)

Knowledge benchmarks
BenchmarkGPT-5.1o3
GPQA Diamond87.6%81.8%
Humanity's Last Exam23.7%20.3%
SimpleQA Verified48%49.4%
MMLU-Pro57.9%85.9%
GPQA (HELM)44.2%75.3%
LMArena Expert14701402
Confabulations—14.4%
Vectara Hallucination Rate10.9%—

Multimodal GPT-5.1 leads

GPT-5.1: 44.8 (#19), o3: 41.4 (#36)

Multimodal benchmarks
BenchmarkGPT-5.1o3
LMArena Vision12501214
VPCT58.7%52%
GeoBench—74%
LMArena Document1403—

Multilingual GPT-5.1 leads

GPT-5.1: 53.8 (#56), o3: 51.7 (#105)

Multilingual benchmarks
BenchmarkGPT-5.1o3
LMArena Non-English14311401
LMArena Chinese14951437
LMArena French14501430
LMArena German14381420
LMArena Japanese14531403
LMArena Korean14011370
LMArena Russian14351406
LMArena Spanish14331395

Instruction Following GPT-5.1 leads

GPT-5.1: 83.9 (#1), o3: 72.8 (#127)

Instruction Following benchmarks
BenchmarkGPT-5.1o3
IFEval93.5%86.9%
LMArena Instruction Following14431368
LiveBench Instruction Following93.3%—

Long Context o3 leads

GPT-5.1: 47.6 (#14), o3: 53.3 (#6)

Long Context benchmarks
BenchmarkGPT-5.1o3
CL-bench23.7%17.8%
LMArena Longer Query14471372
Fiction.LiveBench—88.9%
CL-bench Life17.3%—

Writing & Preference GPT-5.1 leads

GPT-5.1: 64.5 (#55), o3: 63.5 (#64)

Writing & Preference benchmarks
BenchmarkGPT-5.1o3
LMArena Text14431410
LMArena Creative Writing14271359
WildBench86.3%86.1%
LMArena Multi-Turn14501405
Short-Story Creative Writing—83.9%
EQ-Bench Creative Writing—1676
LiveBench Language80.2%—

Frequently asked questions

Is GPT-5.1 better than o3?

GPT-5.1 is the stronger model overall, scoring 49.0 to 47.5 on the Noometry Index.

Which is cheaper, GPT-5.1 or o3?

GPT-5.1 is cheaper. It lists at $1.25 per million input tokens and $10 per million output tokens; o3 lists at $2 and $8.

Is GPT-5.1 or o3 better for coding?

They score almost the same on coding (46.4 vs 46.8); test both on your own repository before choosing.

Which has the bigger context window?

GPT-5.1 does, with 400K tokens against 200K.

How many benchmarks do GPT-5.1 and o3 share?

49 benchmarks have published results for both models. GPT-5.1 has 63 scored results on Noometry and o3 has 63.

Related comparisons

Go deeper