Model comparison

Claude 3.7 Sonnet vs o3

o3 is the stronger model overall, scoring 47.5 to 39.5 on the Noometry Index.

Last verified . 48 shared benchmarks.

Claude 3.7 Sonnet Anthropic

39.5

Rank #164 Confirmed

o3 OpenAI

47.5

Rank #61 Confirmed

Summary

  • They share 48 benchmarks with published results for both. Claude 3.7 Sonnet scores higher in 1 category and o3 in 9 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where o3 leads 54.6 to 39.8.
  • The biggest single-benchmark swing is Omni-MATH: 33% for Claude 3.7 Sonnet and 71.4% for o3.

Side by side

Claude 3.7 Sonnet and o3 specifications
Claude 3.7 Sonneto3
ProviderAnthropicOpenAI
Noometry Index39.547.5
Released2025-02-242025-04-16
WeightsProprietaryProprietary
Context window—200K
Max output—100K
Input $ / M tokens—$2
Output $ / M tokens—$8
Results tracked5863

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding o3 leads

Claude 3.7 Sonnet: 40.6 (#136), o3: 46.8 (#64)

Coding benchmarks
BenchmarkClaude 3.7 Sonneto3
SWE-bench Verified61%62.3%
SWE-bench Verified (bash only)52.8%58.4%
Aider Polyglot64.9%81.3%
GSO3.8%8.8%
LMArena Coding13611408
CadEval54%74%
WeirdML—52.4%
LiveBench Coding74.5%—
ALE-Bench—933.55

Agentic & Tool Use Too close to call

Claude 3.7 Sonnet: 34.1 (#50), o3: 34.5 (#44)

Agentic & Tool Use benchmarks
BenchmarkClaude 3.7 Sonneto3
DeepResearch Bench43.6%45.2%
OSWorld35.8%23%
METR Time Horizons60%65.4%
Berkeley Function Calling Leaderboard—63%
GDPval—30.8%
TheAgentCompany30.9%—
Cybench20%—
LMArena Search—1144

Reasoning o3 leads

Claude 3.7 Sonnet: 18.6 (#277), o3: 32.0 (#78)

Reasoning benchmarks
BenchmarkClaude 3.7 Sonneto3
ARC-AGI-20.9%6.5%
SimpleBench46.4%53.1%
ARC-AGI-128.6%60.8%
EnigmaEval4.2%13.1%
LMArena Hard Prompts13331402
Epoch Capabilities Index141.16146.86
ForecastBench61.862.5
Kagi LLM Benchmark—67.6%
CritPt—1.4%
Chess Puzzles—38%
LiveBench Reasoning87.8%—
Mystery Game Puzzles—29%
DTBench—84.8%
LiveBench Data Analysis74%—
LMCA—39.7%
LiveBench76.1%—

Math o3 leads

Claude 3.7 Sonnet: 37.5 (#153), o3: 50.2 (#58)

Math benchmarks
BenchmarkClaude 3.7 Sonneto3
OTIS Mock AIME 2024-202557.8%84.4%
Omni-MATH33%71.4%
LMArena Math13371426
MATH Level 591.2%97.8%
FrontierMath (Feb 2025 set)4.1%18.7%
FrontierMath (Tiers 1-3)—33.3%
LiveBench Math79%—
FrontierMath Tier 4 (v1)—2.1%

Knowledge o3 leads

Claude 3.7 Sonnet: 39.8 (#130), o3: 54.6 (#52)

Knowledge benchmarks
BenchmarkClaude 3.7 Sonneto3
GPQA Diamond79.7%81.8%
Humanity's Last Exam8%20.3%
MMLU-Pro78.4%85.9%
Confabulations14.7%14.4%
GPQA (HELM)60.8%75.3%
LMArena Expert13211402
SimpleQA Verified—49.4%

Multimodal o3 leads

Claude 3.7 Sonnet: 33.7 (#95), o3: 41.4 (#36)

Multimodal benchmarks
BenchmarkClaude 3.7 Sonneto3
LMArena Vision11691214
GeoBench68%74%
VPCT39%52%
SpatialViz-Bench33.9%—

Multilingual o3 leads

Claude 3.7 Sonnet: 44.1 (#179), o3: 51.7 (#105)

Multilingual benchmarks
BenchmarkClaude 3.7 Sonneto3
LMArena Non-English12961401
LMArena Chinese12991437
LMArena French13031430
LMArena German13011420
LMArena Japanese12671403
LMArena Korean12491370
LMArena Russian13111406
LMArena Spanish12981395

Instruction Following Too close to call

Claude 3.7 Sonnet: 72.9 (#125), o3: 72.8 (#127)

Instruction Following benchmarks
BenchmarkClaude 3.7 Sonneto3
IFEval83.4%86.9%
LMArena Instruction Following13521368
LiveBench Instruction Following81.3%—

Long Context o3 leads

Claude 3.7 Sonnet: 50.3 (#10), o3: 53.3 (#6)

Long Context benchmarks
BenchmarkClaude 3.7 Sonneto3
Fiction.LiveBench83.3%88.9%
LMArena Longer Query13731372
CL-bench—17.8%

Writing & Preference o3 leads

Claude 3.7 Sonnet: 54.4 (#150), o3: 63.5 (#64)

Writing & Preference benchmarks
BenchmarkClaude 3.7 Sonneto3
LMArena Text13141410
LMArena Creative Writing13321359
Short-Story Creative Writing81.1%83.9%
EQ-Bench Creative Writing14121676
WildBench81.4%86.1%
LMArena Multi-Turn13391405
LiveBench Language59.9%—

Frequently asked questions

Is Claude 3.7 Sonnet better than o3?

o3 is the stronger model overall, scoring 47.5 to 39.5 on the Noometry Index.

Is Claude 3.7 Sonnet or o3 better for coding?

o3 scores higher on coding benchmarks: 46.8 versus 40.6 in the Noometry coding category.

How many benchmarks do Claude 3.7 Sonnet and o3 share?

48 benchmarks have published results for both models. Claude 3.7 Sonnet has 58 scored results on Noometry and o3 has 63.

Related comparisons

Go deeper