Model comparison

Claude Sonnet 4.6 vs Claude Sonnet 5

Claude Sonnet 5 is the stronger model overall, scoring 54.6 to 50.3 on the Noometry Index.

Last verified . 44 shared benchmarks.

Claude Sonnet 4.6 Anthropic

50.3

Rank #50 Confirmed

Claude Sonnet 5 Anthropic

54.6

Rank #29 Confirmed

Summary

  • They share 44 benchmarks with published results for both. Claude Sonnet 4.6 scores higher in 4 categories and Claude Sonnet 5 in 6 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in math, where Claude Sonnet 5 leads 66.2 to 52.9.
  • The biggest single-benchmark swing is ProofBench: 45% for Claude Sonnet 4.6 and 77% for Claude Sonnet 5.
  • Claude Sonnet 5 is cheaper at $2 / $10 per million input/output tokens, against $3 / $15 for Claude Sonnet 4.6.

Side by side

Claude Sonnet 4.6 and Claude Sonnet 5 specifications
Claude Sonnet 4.6Claude Sonnet 5
ProviderAnthropicAnthropic
Noometry Index50.354.6
Released2026-02-172026-06-29
WeightsProprietaryProprietary
Context window1M1M
Max output128K128K
Input $ / M tokens$3$2
Output $ / M tokens$15$10
Results tracked5751

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Claude Sonnet 5 leads

Claude Sonnet 4.6: 46.3 (#67), Claude Sonnet 5: 55.5 (#26)

Coding benchmarks
BenchmarkClaude Sonnet 4.6Claude Sonnet 5
DeepSWE29.9%53.8%
FrontierCode24.3%42.7%
LMArena WebDev15221541
SciCode46.8%54.3%
WeirdML66.1%68.8%
LMArena Coding15041483
ALE-Bench1,3271,463
SWE-bench Verified75.2%—
CursorBench—34.1%
GSO—37.3%

Agentic & Tool Use Claude Sonnet 5 leads

Claude Sonnet 4.6: 39.1 (#28), Claude Sonnet 5: 42.8 (#18)

Agentic & Tool Use benchmarks
BenchmarkClaude Sonnet 4.6Claude Sonnet 5
APEX-Agents43%54.5%
GBAEval48.8%65.3%
LMArena Search12211194
Vending-Bench 27,2046,378
Terminal-Bench53.4%—
OSWorld 2.09.3%—
DeepResearch Bench54.9%—
OSWorld72.1%—
ExploitBench23.6%—
GDP.pdf18%—

Reasoning Claude Sonnet 5 leads

Claude Sonnet 4.6: 46.1 (#45), Claude Sonnet 5: 49.1 (#39)

Reasoning benchmarks
BenchmarkClaude Sonnet 4.6Claude Sonnet 5
NYT Connections (extended)80.9%75.1%
CritPt3.1%16.9%
Chess Puzzles13%35%
LMArena Hard Prompts14841461
Mystery Game Puzzles16%35%
DTBench89.9%92.5%
LMCA46.5%50%
Epoch Capabilities Index152.24156.21
ForecastBench6261.1
ARC-AGI-260.4%—
SimpleBench—60.6%
ARC-AGI-186.5%—
Thematic Generalization76.3%—
Surface Evolver Bench—60%
Bench to the Future 3—0.14

Math Claude Sonnet 5 leads

Claude Sonnet 4.6: 52.9 (#49), Claude Sonnet 5: 66.2 (#27)

Math benchmarks
BenchmarkClaude Sonnet 4.6Claude Sonnet 5
OTIS Mock AIME 2024-202585.8%94.7%
ProofBench45%77%
LMArena Math14621467
FrontierMath (Tiers 1-3)—65.6%
FrontierMath Tier 4—29.3%
FrontierMath (Feb 2025 set)32.4%—
FrontierMath Tier 4 (v1)8.3%—

Knowledge Claude Sonnet 5 leads

Claude Sonnet 4.6: 51.7 (#65), Claude Sonnet 5: 55.6 (#47)

Knowledge benchmarks
BenchmarkClaude Sonnet 4.6Claude Sonnet 5
GPQA Diamond87.4%90.5%
SimpleQA Verified35.5%33.7%
LMArena Expert15001490
Vectara Hallucination Rate10.6%—

Multimodal Claude Sonnet 5 leads

Claude Sonnet 4.6: 38.0 (#68), Claude Sonnet 5: 42.4 (#31)

Multimodal benchmarks
BenchmarkClaude Sonnet 4.6Claude Sonnet 5
LMArena Vision12831274
Blueprint-Bench 26.7%24.9%
LMArena Document14821466

Multilingual Too close to call

Claude Sonnet 4.6: 54.4 (#41), Claude Sonnet 5: 53.8 (#55)

Multilingual benchmarks
BenchmarkClaude Sonnet 4.6Claude Sonnet 5
LMArena Non-English14401431
LMArena Chinese14911477
LMArena French14651460
LMArena German14281440
LMArena Japanese14201422
LMArena Korean14111411
LMArena Russian14401451
LMArena Spanish14641437

Instruction Following Claude Sonnet 4.6 leads

Claude Sonnet 4.6: 77.4 (#25), Claude Sonnet 5: 76.3 (#41)

Instruction Following benchmarks
BenchmarkClaude Sonnet 4.6Claude Sonnet 5
LMArena Instruction Following14751452

Long Context Too close to call

Claude Sonnet 4.6: 45.3 (#44), Claude Sonnet 5: 44.8 (#55)

Long Context benchmarks
BenchmarkClaude Sonnet 4.6Claude Sonnet 5
LMArena Longer Query14791463

Writing & Preference Too close to call

Claude Sonnet 4.6: 70.2 (#22), Claude Sonnet 5: 69.2 (#25)

Writing & Preference benchmarks
BenchmarkClaude Sonnet 4.6Claude Sonnet 5
LMArena Text14581442
LMArena Creative Writing14351416
EQ-Bench Creative Writing18101794
EQ-Bench 412071236
LMArena Multi-Turn14641454

Frequently asked questions

Is Claude Sonnet 4.6 better than Claude Sonnet 5?

Claude Sonnet 5 is the stronger model overall, scoring 54.6 to 50.3 on the Noometry Index.

Which is cheaper, Claude Sonnet 4.6 or Claude Sonnet 5?

Claude Sonnet 5 is cheaper. It lists at $2 per million input tokens and $10 per million output tokens; Claude Sonnet 4.6 lists at $3 and $15.

Is Claude Sonnet 4.6 or Claude Sonnet 5 better for coding?

Claude Sonnet 5 scores higher on coding benchmarks: 55.5 versus 46.3 in the Noometry coding category.

Which has the bigger context window?

Both accept 1M tokens.

How many benchmarks do Claude Sonnet 4.6 and Claude Sonnet 5 share?

44 benchmarks have published results for both models. Claude Sonnet 4.6 has 57 scored results on Noometry and Claude Sonnet 5 has 51.

Related comparisons

Go deeper