Model comparison

Claude 3.7 Sonnet vs Claude Sonnet 4

Claude Sonnet 4 is the stronger model overall, scoring 40.8 to 39.5 on the Noometry Index.

Last verified . 48 shared benchmarks.

Claude 3.7 Sonnet Anthropic

39.5

Rank #164 Confirmed

Claude Sonnet 4 Anthropic

40.8

Rank #145 Confirmed

Summary

  • They share 48 benchmarks with published results for both. Claude 3.7 Sonnet scores higher in 3 categories and Claude Sonnet 4 in 7 categories; 10 gaps are clear of the uncertainty.
  • The widest gap is in long context, where Claude 3.7 Sonnet leads 50.3 to 33.7.
  • The biggest single-benchmark swing is Fiction.LiveBench: 83.3% for Claude 3.7 Sonnet and 46.9% for Claude Sonnet 4.

Side by side

Claude 3.7 Sonnet and Claude Sonnet 4 specifications
Claude 3.7 SonnetClaude Sonnet 4
ProviderAnthropicAnthropic
Noometry Index39.540.8
Released2025-02-242025-05-22
WeightsProprietaryProprietary
Context window—200K
Max output—64K
Input $ / M tokens—$3
Output $ / M tokens—$15
Results tracked5858

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Claude Sonnet 4 leads

Claude 3.7 Sonnet: 40.6 (#136), Claude Sonnet 4: 43.5 (#88)

Coding benchmarks
BenchmarkClaude 3.7 SonnetClaude Sonnet 4
SWE-bench Verified (bash only)52.8%64.9%
Aider Polyglot64.9%61.3%
GSO3.8%4.9%
LMArena Coding13611414
SWE-bench Verified61%—
SciCode—40%
WeirdML—46.1%
LiveBench Coding74.5%—
CadEval54%—
ALE-Bench—655.35

Agentic & Tool Use Claude Sonnet 4 leads

Claude 3.7 Sonnet: 34.1 (#50), Claude Sonnet 4: 38.5 (#31)

Agentic & Tool Use benchmarks
BenchmarkClaude 3.7 SonnetClaude Sonnet 4
TheAgentCompany30.9%33.1%
Cybench20%35%
DeepResearch Bench43.6%46.6%
OSWorld35.8%43.9%
METR Time Horizons60%62%

Reasoning Claude Sonnet 4 leads

Claude 3.7 Sonnet: 18.6 (#277), Claude Sonnet 4: 22.9 (#187)

Reasoning benchmarks
BenchmarkClaude 3.7 SonnetClaude Sonnet 4
ARC-AGI-20.9%5.9%
SimpleBench46.4%45.5%
ARC-AGI-128.6%40%
EnigmaEval4.2%3.1%
LMArena Hard Prompts13331372
Epoch Capabilities Index141.16141.69
ForecastBench61.860.2
Kagi LLM Benchmark—73%
CritPt—0.3%
LiveBench Reasoning87.8%—
DTBench—77.1%
LiveBench Data Analysis74%—
LMCA—29%
LiveBench76.1%—

Math Claude Sonnet 4 leads

Claude 3.7 Sonnet: 37.5 (#153), Claude Sonnet 4: 43.3 (#80)

Math benchmarks
BenchmarkClaude 3.7 SonnetClaude Sonnet 4
OTIS Mock AIME 2024-202557.8%71.1%
Omni-MATH33%60.2%
LMArena Math13371375
MATH Level 591.2%84.4%
FrontierMath (Feb 2025 set)4.1%4.1%
LiveBench Math79%—
FrontierMath Tier 4 (v1)—0%

Knowledge Claude Sonnet 4 leads

Claude 3.7 Sonnet: 39.8 (#130), Claude Sonnet 4: 41.8 (#108)

Knowledge benchmarks
BenchmarkClaude 3.7 SonnetClaude Sonnet 4
GPQA Diamond79.7%79.2%
Humanity's Last Exam8%7.8%
MMLU-Pro78.4%84.3%
Confabulations14.7%13.2%
GPQA (HELM)60.8%70.6%
LMArena Expert13211372
Vectara Hallucination Rate—10.3%

Multimodal Claude 3.7 Sonnet leads

Claude 3.7 Sonnet: 33.7 (#95), Claude Sonnet 4: 26.2 (#121)

Multimodal benchmarks
BenchmarkClaude 3.7 SonnetClaude Sonnet 4
LMArena Vision11691191
GeoBench68%37%
VPCT39%34%
MindCube—44.8%
SpatialViz-Bench33.9%—

Multilingual Claude Sonnet 4 leads

Claude 3.7 Sonnet: 44.1 (#179), Claude Sonnet 4: 46.7 (#156)

Multilingual benchmarks
BenchmarkClaude 3.7 SonnetClaude Sonnet 4
LMArena Non-English12961333
LMArena Chinese12991350
LMArena French13031363
LMArena German13011331
LMArena Japanese12671302
LMArena Korean12491291
LMArena Russian13111355
LMArena Spanish12981357

Instruction Following Claude 3.7 Sonnet leads

Claude 3.7 Sonnet: 72.9 (#125), Claude Sonnet 4: 71.7 (#145)

Instruction Following benchmarks
BenchmarkClaude 3.7 SonnetClaude Sonnet 4
IFEval83.4%84%
LMArena Instruction Following13521376
LiveBench Instruction Following81.3%—

Long Context Claude 3.7 Sonnet leads

Claude 3.7 Sonnet: 50.3 (#10), Claude Sonnet 4: 33.7 (#259)

Long Context benchmarks
BenchmarkClaude 3.7 SonnetClaude Sonnet 4
Fiction.LiveBench83.3%46.9%
LMArena Longer Query13731398

Writing & Preference Claude Sonnet 4 leads

Claude 3.7 Sonnet: 54.4 (#150), Claude Sonnet 4: 57.1 (#132)

Writing & Preference benchmarks
BenchmarkClaude 3.7 SonnetClaude Sonnet 4
LMArena Text13141351
LMArena Creative Writing13321345
Short-Story Creative Writing81.1%81.4%
EQ-Bench Creative Writing14121483
WildBench81.4%83.8%
LMArena Multi-Turn13391376
LiveBench Language59.9%—

Frequently asked questions

Is Claude 3.7 Sonnet better than Claude Sonnet 4?

Claude Sonnet 4 is the stronger model overall, scoring 40.8 to 39.5 on the Noometry Index.

Is Claude 3.7 Sonnet or Claude Sonnet 4 better for coding?

Claude Sonnet 4 scores higher on coding benchmarks: 43.5 versus 40.6 in the Noometry coding category.

How many benchmarks do Claude 3.7 Sonnet and Claude Sonnet 4 share?

48 benchmarks have published results for both models. Claude 3.7 Sonnet has 58 scored results on Noometry and Claude Sonnet 4 has 58.

Related comparisons

Go deeper