Model comparison

Claude Opus 4 vs Claude Opus 4.8

Claude Opus 4.8 is the stronger model overall, scoring 60.7 to 43.1 on the Noometry Index.

Last verified . 37 shared benchmarks.

Claude Opus 4 Anthropic

43.1

Rank #100 Confirmed

Claude Opus 4.8 Anthropic

60.7

Rank #13 Confirmed

Summary

  • They share 37 benchmarks with published results for both. Claude Opus 4 scores higher in 0 categories and Claude Opus 4.8 in 10 categories; 9 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Claude Opus 4.8 leads 64.7 to 27.3.
  • The biggest single-benchmark swing is ARC-AGI-2: 8.6% for Claude Opus 4 and 72.1% for Claude Opus 4.8.
  • Claude Opus 4.8 is cheaper at $5 / $25 per million input/output tokens, against $15 / $75 for Claude Opus 4.
  • Claude Opus 4.8 accepts more context: 1M tokens versus 200K.

Side by side

Claude Opus 4 and Claude Opus 4.8 specifications
Claude Opus 4Claude Opus 4.8
ProviderAnthropicAnthropic
Noometry Index43.160.7
Released2025-05-222026-05-28
WeightsProprietaryProprietary
Context window200K1M
Max output32K128K
Input $ / M tokens$15$5
Output $ / M tokens$75$25
Results tracked5665

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Claude Opus 4.8 leads

Claude Opus 4: 47.2 (#62), Claude Opus 4.8: 59.9 (#12)

Coding benchmarks
BenchmarkClaude Opus 4Claude Opus 4.8
GSO6.9%47.1%
WeirdML43.7%82.9%
LMArena Coding14421490
SWE-bench Verified70.7%—
DeepSWE—59%
FrontierCode—46.5%
SWE-bench Verified (bash only)67.6%—
Aider Polyglot72%—
LMArena WebDev—1556
SciCode—53.5%
ALE-Bench—1,564
AlgoTune1.33—

Agentic & Tool Use Claude Opus 4.8 leads

Claude Opus 4: 34.8 (#42), Claude Opus 4.8: 47.6 (#11)

Agentic & Tool Use benchmarks
BenchmarkClaude Opus 4Claude Opus 4.8
DeepResearch Bench46.8%50.2%
LMArena Search11271204
APEX-Agents—48.9%
OSWorld 2.0—20.6%
Remote Labor Index—8.3%
τ²-bench Banking—39.7%
Cybench38%—
PostTrainBench—33.8%
GBAEval—70.9%
GDP.pdf—24%
METR Time Horizons63.9%—
Vending-Bench 2—5,787

Reasoning Claude Opus 4.8 leads

Claude Opus 4: 27.3 (#121), Claude Opus 4.8: 64.7 (#16)

Reasoning benchmarks
BenchmarkClaude Opus 4Claude Opus 4.8
ARC-AGI-28.6%72.1%
SimpleBench58.8%64.8%
Kagi LLM Benchmark74.3%88.8%
ARC-AGI-135.7%92.5%
CritPt0.3%20.9%
EnigmaEval5.6%23.5%
LMArena Hard Prompts13991482
DTBench81.6%94.9%
LMCA37.4%57.5%
Epoch Capabilities Index142.67158.21
ForecastBench61.159.9
NYT Connections (extended)—91.1%
Chess Puzzles—34%
EBR-Bench—28.6%
Mystery Game Puzzles—36%
Surface Evolver Bench—87.5%
Bench to the Future 3—0.14

Math Claude Opus 4.8 leads

Claude Opus 4: 42.0 (#86), Claude Opus 4.8: 78.4 (#13)

Knowledge Claude Opus 4.8 leads

Claude Opus 4: 44.0 (#88), Claude Opus 4.8: 61.3 (#29)

Knowledge benchmarks
BenchmarkClaude Opus 4Claude Opus 4.8
GPQA Diamond76.3%91%
LMArena Expert13861502
Humanity's Last Exam10.7%—
SimpleQA Verified—53%
MMLU-Pro87.5%—
Confabulations15.9%—
Vectara Hallucination Rate12%—
GPQA (HELM)70.8%—

Multimodal Claude Opus 4.8 leads

Claude Opus 4: 31.5 (#106), Claude Opus 4.8: 42.9 (#26)

Multimodal benchmarks
BenchmarkClaude Opus 4Claude Opus 4.8
LMArena Vision11921294
GeoBench49%—
VPCT38%—
Blueprint-Bench 2—14.5%
Furniture Assembly—42.5%
LMArena Document—1475

Multilingual Claude Opus 4.8 leads

Claude Opus 4: 48.8 (#138), Claude Opus 4.8: 55.2 (#33)

Multilingual benchmarks
BenchmarkClaude Opus 4Claude Opus 4.8
LMArena Non-English13621450
LMArena Chinese13861507
LMArena French13721481
LMArena German13911472
LMArena Japanese13311440
LMArena Korean13211432
LMArena Russian13921474
LMArena Spanish13891466

Instruction Following Too close to call

Claude Opus 4: 77.1 (#28), Claude Opus 4.8: 77.4 (#24)

Instruction Following benchmarks
BenchmarkClaude Opus 4Claude Opus 4.8
LMArena Instruction Following14061476
IFEval91.8%—

Long Context Claude Opus 4.8 leads

Claude Opus 4: 39.6 (#172), Claude Opus 4.8: 45.4 (#35)

Long Context benchmarks
BenchmarkClaude Opus 4Claude Opus 4.8
LMArena Longer Query14221483
Fiction.LiveBench61.1%—

Writing & Preference Claude Opus 4.8 leads

Claude Opus 4: 61.2 (#89), Claude Opus 4.8: 72.0 (#16)

Writing & Preference benchmarks
BenchmarkClaude Opus 4Claude Opus 4.8
LMArena Text13771461
LMArena Creative Writing13871454
EQ-Bench Creative Writing15801840
LMArena Multi-Turn13961476
Short-Story Creative Writing83.6%—
WildBench85.2%—
EQ-Bench 4—1281

Frequently asked questions

Is Claude Opus 4 better than Claude Opus 4.8?

Claude Opus 4.8 is the stronger model overall, scoring 60.7 to 43.1 on the Noometry Index.

Which is cheaper, Claude Opus 4 or Claude Opus 4.8?

Claude Opus 4.8 is cheaper. It lists at $5 per million input tokens and $25 per million output tokens; Claude Opus 4 lists at $15 and $75.

Is Claude Opus 4 or Claude Opus 4.8 better for coding?

Claude Opus 4.8 scores higher on coding benchmarks: 59.9 versus 47.2 in the Noometry coding category.

Which has the bigger context window?

Claude Opus 4.8 does, with 1M tokens against 200K.

How many benchmarks do Claude Opus 4 and Claude Opus 4.8 share?

37 benchmarks have published results for both models. Claude Opus 4 has 56 scored results on Noometry and Claude Opus 4.8 has 65.

Related comparisons

Go deeper