Model comparison

Claude 2 vs Claude Opus 5.5

Claude Opus 5.5 is the stronger model overall, scoring 68.6 to 25.0 on the Noometry Index.

Last verified . 4 shared benchmarks.

Claude 2 Anthropic

25.0

Rank #346 Reported

Claude Opus 5.5 Anthropic

68.6

Rank #3 Confirmed

Summary

  • They share 4 benchmarks with published results for both. Claude 2 scores higher in 0 categories and Claude Opus 5.5 in 3 categories; 3 gaps are clear of the uncertainty.
  • The widest gap is in math, where Claude Opus 5.5 leads 91.8 to 9.3.
  • The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 2.5% for Claude 2 and 100% for Claude Opus 5.5.

Side by side

Claude 2 and Claude Opus 5.5 specifications
Claude 2Claude Opus 5.5
ProviderAnthropicAnthropic
Noometry Index25.068.6
Released2023-07-112026-09-22
WeightsProprietaryProprietary
Context window—1M
Max output—128K
Input $ / M tokens—$4
Output $ / M tokens—$20
Results tracked844

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Not comparable

Claude 2: —, Claude Opus 5.5: 71.9 (#3)

Coding benchmarks
BenchmarkClaude 2Claude Opus 5.5
FrontierCode—54.6%
CursorBench—57.8%
LMArena WebDev—1813
FrontierSWE—62.3%
SciCode—66.9%
LMArena Coding—1547
MirrorCode—77.4%
ALE-Bench—2,147
HumanEval+61.6%—

Agentic & Tool Use Not comparable

Claude 2: —, Claude Opus 5.5: 45.3 (#15)

Agentic & Tool Use benchmarks
BenchmarkClaude 2Claude Opus 5.5
APEX-Agents—73.5%
GDP.pdf—30.6%
Vending-Bench 2—9,235

Reasoning Claude Opus 5.5 leads

Claude 2: 21.7 (#216), Claude Opus 5.5: 80.2 (#3)

Reasoning benchmarks
BenchmarkClaude 2Claude Opus 5.5
DTBench51.9%98.9%
Epoch Capabilities Index120.13167.33
ARC-AGI-2—93.3%
NYT Connections (extended)—88.5%
ARC-AGI-1—98.5%
CritPt—31.7%
EBR-Bench—71.4%
LMArena Hard Prompts—1535
Mystery Game Puzzles—71%
LMCA—68.2%

Math Claude Opus 5.5 leads

Claude 2: 9.3 (#320), Claude Opus 5.5: 91.8 (#3)

Math benchmarks
BenchmarkClaude 2Claude Opus 5.5
OTIS Mock AIME 2024-20252.5%100%
FrontierMath (Tiers 1-3)—91.2%
FrontierMath Tier 4—95%
ProofBench—100%
LMArena Math—1506
MATH Level 511.7%—
FrontierMath Erdős—2.9%

Knowledge Claude Opus 5.5 leads

Claude 2: 16.9 (#287), Claude Opus 5.5: 66.4 (#10)

Knowledge benchmarks
BenchmarkClaude 2Claude Opus 5.5
GPQA Diamond34.7%90.6%
SimpleQA Verified—72.2%
LMArena Expert—1547
MMLU78.5%—
TriviaQA87.5%—

Multimodal Not comparable

Claude 2: —, Claude Opus 5.5: 57.8 (#1)

Multimodal benchmarks
BenchmarkClaude 2Claude Opus 5.5
LMArena Vision—1321
Blueprint-Bench 2—51.2%
Furniture Assembly—83.3%

Multilingual Not comparable

Claude 2: —, Claude Opus 5.5: 59.1 (#2)

Multilingual benchmarks
BenchmarkClaude 2Claude Opus 5.5
LMArena Non-English—1507
LMArena Chinese—1588
LMArena French—1514
LMArena Russian—1520
LMArena Spanish—1507

Instruction Following Not comparable

Claude 2: —, Claude Opus 5.5: 80.0 (#3)

Instruction Following benchmarks
BenchmarkClaude 2Claude Opus 5.5
LMArena Instruction Following—1537

Long Context Not comparable

Claude 2: —, Claude Opus 5.5: 47.1 (#19)

Long Context benchmarks
BenchmarkClaude 2Claude Opus 5.5
LMArena Longer Query—1532

Writing & Preference Not comparable

Claude 2: —, Claude Opus 5.5: 78.2 (#3)

Writing & Preference benchmarks
BenchmarkClaude 2Claude Opus 5.5
LMArena Text—1515
LMArena Creative Writing—1533
EQ-Bench Creative Writing—2050
LMArena Multi-Turn—1499

Frequently asked questions

Is Claude 2 better than Claude Opus 5.5?

Claude Opus 5.5 is the stronger model overall, scoring 68.6 to 25.0 on the Noometry Index.

How many benchmarks do Claude 2 and Claude Opus 5.5 share?

4 benchmarks have published results for both models. Claude 2 has 8 scored results on Noometry and Claude Opus 5.5 has 44.

Related comparisons

Go deeper