Model comparison

Claude 3.7 Sonnet vs o1

o1 is the stronger model overall, scoring 40.9 to 39.5 on the Noometry Index.

Last verified . 44 shared benchmarks.

Claude 3.7 Sonnet Anthropic

39.5

Rank #164 Confirmed

o1 OpenAI

40.9

Rank #143 Confirmed

Summary

  • They share 44 benchmarks with published results for both. Claude 3.7 Sonnet scores higher in 2 categories and o1 in 8 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in agentic & tool use, where Claude 3.7 Sonnet leads 34.1 to 24.6.
  • The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 57.8% for Claude 3.7 Sonnet and 73.3% for o1.

Side by side

Claude 3.7 Sonnet and o1 specifications
Claude 3.7 Sonneto1
ProviderAnthropicOpenAI
Noometry Index39.540.9
Released2025-02-242024-09-12
WeightsProprietaryProprietary
Context window—200K
Max output—100K
Input $ / M tokens—$15
Output $ / M tokens—$60
Results tracked5852

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding o1 leads

Claude 3.7 Sonnet: 40.6 (#136), o1: 46.1 (#70)

Coding benchmarks
BenchmarkClaude 3.7 Sonneto1
Aider Polyglot64.9%61.7%
LiveBench Coding74.5%69.7%
LMArena Coding13611367
CadEval54%56%
SWE-bench Verified61%—
SWE-bench Verified (bash only)52.8%—
GSO3.8%—
WeirdML—47.6%
HumanEval+—89%
MBPP+—80.2%

Agentic & Tool Use Claude 3.7 Sonnet leads

Claude 3.7 Sonnet: 34.1 (#50), o1: 24.6 (#117)

Agentic & Tool Use benchmarks
BenchmarkClaude 3.7 Sonneto1
Cybench20%10%
METR Time Horizons60%51.1%
TheAgentCompany30.9%—
DeepResearch Bench43.6%—
OSWorld35.8%—

Reasoning o1 leads

Claude 3.7 Sonnet: 18.6 (#277), o1: 27.9 (#111)

Reasoning benchmarks
BenchmarkClaude 3.7 Sonneto1
SimpleBench46.4%41.7%
ARC-AGI-128.6%30.7%
EnigmaEval4.2%5.7%
LiveBench Reasoning87.8%91.6%
LMArena Hard Prompts13331371
LiveBench Data Analysis74%65.5%
Epoch Capabilities Index141.16141.91
LiveBench76.1%75.7%
ARC-AGI-20.9%—
Chess Puzzles—15%
DTBench—74.7%
LMCA—22.3%
ForecastBench61.8—

Math Claude 3.7 Sonnet leads

Claude 3.7 Sonnet: 37.5 (#153), o1: 36.1 (#175)

Math benchmarks
BenchmarkClaude 3.7 Sonneto1
OTIS Mock AIME 2024-202557.8%73.3%
LiveBench Math79%80.3%
LMArena Math13371388
MATH Level 591.2%94.7%
FrontierMath (Feb 2025 set)4.1%9.3%
FrontierMath (Tiers 1-3)—14.7%
Omni-MATH33%—

Knowledge o1 leads

Claude 3.7 Sonnet: 39.8 (#130), o1: 41.5 (#110)

Knowledge benchmarks
BenchmarkClaude 3.7 Sonneto1
GPQA Diamond79.7%76.8%
Humanity's Last Exam8%8%
Confabulations14.7%11.7%
LMArena Expert13211361
SimpleQA Verified—41.1%
MMLU-Pro78.4%—
GPQA (HELM)60.8%—

Multimodal Too close to call

Claude 3.7 Sonnet: 33.7 (#95), o1: 34.2 (#93)

Multimodal benchmarks
BenchmarkClaude 3.7 Sonneto1
LMArena Vision11691168
GeoBench68%80%
VPCT39%37%
SpatialViz-Bench33.9%41.4%

Multilingual o1 leads

Claude 3.7 Sonnet: 44.1 (#179), o1: 48.6 (#142)

Multilingual benchmarks
BenchmarkClaude 3.7 Sonneto1
LMArena Non-English12961358
LMArena Chinese12991394
LMArena French13031344
LMArena German13011337
LMArena Japanese12671346
LMArena Korean12491396
LMArena Russian13111356
LMArena Spanish12981345

Instruction Following o1 leads

Claude 3.7 Sonnet: 72.9 (#125), o1: 74.8 (#86)

Instruction Following benchmarks
BenchmarkClaude 3.7 Sonneto1
LiveBench Instruction Following81.3%81.5%
LMArena Instruction Following13521367
IFEval83.4%—

Long Context Too close to call

Claude 3.7 Sonnet: 50.3 (#10), o1: 50.3 (#9)

Long Context benchmarks
BenchmarkClaude 3.7 Sonneto1
Fiction.LiveBench83.3%83.3%
LMArena Longer Query13731378

Writing & Preference o1 leads

Claude 3.7 Sonnet: 54.4 (#150), o1: 55.6 (#144)

Writing & Preference benchmarks
BenchmarkClaude 3.7 Sonneto1
LMArena Text13141366
LMArena Creative Writing13321348
Short-Story Creative Writing81.1%70.2%
LMArena Multi-Turn13391369
LiveBench Language59.9%65.4%
EQ-Bench Creative Writing1412—
WildBench81.4%—

Frequently asked questions

Is Claude 3.7 Sonnet better than o1?

o1 is the stronger model overall, scoring 40.9 to 39.5 on the Noometry Index.

Is Claude 3.7 Sonnet or o1 better for coding?

o1 scores higher on coding benchmarks: 46.1 versus 40.6 in the Noometry coding category.

How many benchmarks do Claude 3.7 Sonnet and o1 share?

44 benchmarks have published results for both models. Claude 3.7 Sonnet has 58 scored results on Noometry and o1 has 52.

Related comparisons

Go deeper