Model comparison

Claude 3 Opus vs Claude Opus 5.5

Claude Opus 5.5 is the stronger model overall, scoring 68.6 to 29.5 on the Noometry Index.

Last verified . 21 shared benchmarks.

Claude 3 Opus Anthropic

29.5

Rank #310 Confirmed

Claude Opus 5.5 Anthropic

68.6

Rank #3 Confirmed

Summary

  • They share 21 benchmarks with published results for both. Claude 3 Opus scores higher in 0 categories and Claude Opus 5.5 in 10 categories; 10 gaps are clear of the uncertainty.
  • The widest gap is in math, where Claude Opus 5.5 leads 91.8 to 14.8.
  • The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 4.7% for Claude 3 Opus and 100% for Claude Opus 5.5.

Side by side

Claude 3 Opus and Claude Opus 5.5 specifications
Claude 3 OpusClaude Opus 5.5
ProviderAnthropicAnthropic
Noometry Index29.568.6
Released2024-02-292026-09-22
WeightsProprietaryProprietary
Context window—1M
Max output—128K
Input $ / M tokens—$4
Output $ / M tokens—$20
Results tracked4644

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Claude Opus 5.5 leads

Claude 3 Opus: 32.9 (#267), Claude Opus 5.5: 71.9 (#3)

Coding benchmarks
BenchmarkClaude 3 OpusClaude Opus 5.5
LMArena Coding12641547
FrontierCode—54.6%
CursorBench—57.8%
LMArena WebDev—1813
FrontierSWE—62.3%
SciCode—66.9%
WeirdML19.2%—
BigCodeBench Instruct45.5%—
LiveBench Coding38.6%—
MirrorCode—77.4%
BigCodeBench Complete57.4%—
ALE-Bench—2,147
HumanEval+77.4%—
MBPP+73.3%—

Agentic & Tool Use Claude Opus 5.5 leads

Claude 3 Opus: 24.6 (#116), Claude Opus 5.5: 45.3 (#15)

Agentic & Tool Use benchmarks
BenchmarkClaude 3 OpusClaude Opus 5.5
APEX-Agents—73.5%
Cybench10%—
GDP.pdf—30.6%
METR Time Horizons29.5%—
Vending-Bench 2—9,235

Reasoning Claude Opus 5.5 leads

Claude 3 Opus: 14.6 (#324), Claude Opus 5.5: 80.2 (#3)

Reasoning benchmarks
BenchmarkClaude 3 OpusClaude Opus 5.5
LMArena Hard Prompts12451535
DTBench61.6%98.9%
LMCA17%68.2%
Epoch Capabilities Index126.91167.33
ARC-AGI-2—93.3%
SimpleBench23.5%—
NYT Connections (extended)—88.5%
ARC-AGI-1—98.5%
CritPt—31.7%
Chess Puzzles5%—
EnigmaEval0.8%—
EBR-Bench—71.4%
LiveBench Reasoning40.6%—
Mystery Game Puzzles—71%
LiveBench Data Analysis57.9%—
ForecastBench58.4—
LiveBench49.2%—
WinoGrande88.5%—

Math Claude Opus 5.5 leads

Claude 3 Opus: 14.8 (#299), Claude Opus 5.5: 91.8 (#3)

Math benchmarks
BenchmarkClaude 3 OpusClaude Opus 5.5
OTIS Mock AIME 2024-20254.7%100%
LMArena Math12731506
FrontierMath (Tiers 1-3)—91.2%
FrontierMath Tier 4—95%
ProofBench—100%
LiveBench Math43.6%—
MATH Level 537.5%—
FrontierMath Erdős—2.9%

Knowledge Claude Opus 5.5 leads

Claude 3 Opus: 24.5 (#267), Claude Opus 5.5: 66.4 (#10)

Knowledge benchmarks
BenchmarkClaude 3 OpusClaude Opus 5.5
GPQA Diamond47.2%90.6%
SimpleQA Verified12.6%72.2%
LMArena Expert12231547
Confabulations22.7%—
MMLU84.6%—

Multimodal Claude Opus 5.5 leads

Claude 3 Opus: 27.1 (#116), Claude Opus 5.5: 57.8 (#1)

Multimodal benchmarks
BenchmarkClaude 3 OpusClaude Opus 5.5
LMArena Vision10231321
Blueprint-Bench 2—51.2%
Furniture Assembly—83.3%

Multilingual Claude Opus 5.5 leads

Claude 3 Opus: 41.4 (#207), Claude Opus 5.5: 59.1 (#2)

Multilingual benchmarks
BenchmarkClaude 3 OpusClaude Opus 5.5
LMArena Non-English12581507
LMArena Chinese12481588
LMArena French12751514
LMArena Russian12801520
LMArena Spanish12461507
LMArena German1258—
LMArena Japanese1204—
LMArena Korean1187—

Instruction Following Claude Opus 5.5 leads

Claude 3 Opus: 64.1 (#228), Claude Opus 5.5: 80.0 (#3)

Instruction Following benchmarks
BenchmarkClaude 3 OpusClaude Opus 5.5
LMArena Instruction Following12481537
LiveBench Instruction Following63.9%—

Long Context Claude Opus 5.5 leads

Claude 3 Opus: 38.2 (#202), Claude Opus 5.5: 47.1 (#19)

Long Context benchmarks
BenchmarkClaude 3 OpusClaude Opus 5.5
LMArena Longer Query12591532

Writing & Preference Claude Opus 5.5 leads

Claude 3 Opus: 47.2 (#213), Claude Opus 5.5: 78.2 (#3)

Writing & Preference benchmarks
BenchmarkClaude 3 OpusClaude Opus 5.5
LMArena Text12621515
LMArena Creative Writing12351533
LMArena Multi-Turn12751499
EQ-Bench Creative Writing—2050
LiveBench Language50.4%—

Frequently asked questions

Is Claude 3 Opus better than Claude Opus 5.5?

Claude Opus 5.5 is the stronger model overall, scoring 68.6 to 29.5 on the Noometry Index.

Is Claude 3 Opus or Claude Opus 5.5 better for coding?

Claude Opus 5.5 scores higher on coding benchmarks: 71.9 versus 32.9 in the Noometry coding category.

How many benchmarks do Claude 3 Opus and Claude Opus 5.5 share?

21 benchmarks have published results for both models. Claude 3 Opus has 46 scored results on Noometry and Claude Opus 5.5 has 44.

Related comparisons

Go deeper