Model comparison

Claude Opus 4 vs Claude Opus 4.6

Claude Opus 4.6 is the stronger model overall, scoring 58.2 to 43.1 on the Noometry Index.

Last verified . 43 shared benchmarks.

Claude Opus 4 Anthropic

43.1

Rank #100 Confirmed

Claude Opus 4.6 Anthropic

58.2

Rank #20 Confirmed

Summary

  • They share 43 benchmarks with published results for both. Claude Opus 4 scores higher in 0 categories and Claude Opus 4.6 in 10 categories; 10 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Claude Opus 4.6 leads 57.8 to 27.3.
  • The biggest single-benchmark swing is ARC-AGI-2: 8.6% for Claude Opus 4 and 69.2% for Claude Opus 4.6.
  • Claude Opus 4.6 is cheaper at $5 / $25 per million input/output tokens, against $15 / $75 for Claude Opus 4.
  • Claude Opus 4.6 accepts more context: 1M tokens versus 200K.

Side by side

Claude Opus 4 and Claude Opus 4.6 specifications
Claude Opus 4Claude Opus 4.6
ProviderAnthropicAnthropic
Noometry Index43.158.2
Released2025-05-222026-02-04
WeightsProprietaryProprietary
Context window200K1M
Max output32K128K
Input $ / M tokens$15$5
Output $ / M tokens$75$25
Results tracked5668

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Claude Opus 4.6 leads

Claude Opus 4: 47.2 (#62), Claude Opus 4.6: 57.2 (#20)

Coding benchmarks
BenchmarkClaude Opus 4Claude Opus 4.6
SWE-bench Verified70.7%78.7%
SWE-bench Verified (bash only)67.6%75.6%
GSO6.9%41.2%
WeirdML43.7%78%
LMArena Coding14421536
AlgoTune1.331.47
FrontierCode—26.6%
Aider Polyglot72%—
LMArena WebDev—1547
SWE-bench Multilingual—72%
ALE-Bench—996.5

Agentic & Tool Use Claude Opus 4.6 leads

Claude Opus 4: 34.8 (#42), Claude Opus 4.6: 51.1 (#4)

Agentic & Tool Use benchmarks
BenchmarkClaude Opus 4Claude Opus 4.6
Cybench38%93%
DeepResearch Bench46.8%55.3%
LMArena Search11271253
METR Time Horizons63.9%78.9%
Terminal-Bench—79.8%
APEX-Agents—46.3%
Remote Labor Index—4.2%
τ²-bench Banking—27.3%
GBAEval—44.1%
Vending-Bench 2—8,018

Reasoning Claude Opus 4.6 leads

Claude Opus 4: 27.3 (#121), Claude Opus 4.6: 57.8 (#23)

Reasoning benchmarks
BenchmarkClaude Opus 4Claude Opus 4.6
ARC-AGI-28.6%69.2%
SimpleBench58.8%67.6%
Kagi LLM Benchmark74.3%83.6%
ARC-AGI-135.7%94%
EnigmaEval5.6%7.6%
LMArena Hard Prompts13991527
DTBench81.6%91.2%
LMCA37.4%55.8%
Epoch Capabilities Index142.67155.24
ForecastBench61.160
NYT Connections (extended)—92.1%
CritPt0.3%—
Chess Puzzles—17%
Thematic Generalization—80.6%
EBR-Bench—12.7%
Mystery Game Puzzles—25%

Math Claude Opus 4.6 leads

Claude Opus 4: 42.0 (#86), Claude Opus 4.6: 63.0 (#31)

Knowledge Claude Opus 4.6 leads

Claude Opus 4: 44.0 (#88), Claude Opus 4.6: 61.9 (#26)

Knowledge benchmarks
BenchmarkClaude Opus 4Claude Opus 4.6
GPQA Diamond76.3%90.5%
Humanity's Last Exam10.7%34.4%
Vectara Hallucination Rate12%12.2%
LMArena Expert13861546
SimpleQA Verified—47%
MMLU-Pro87.5%—
Confabulations15.9%—
GPQA (HELM)70.8%—

Multimodal Claude Opus 4.6 leads

Claude Opus 4: 31.5 (#106), Claude Opus 4.6: 37.3 (#74)

Multimodal benchmarks
BenchmarkClaude Opus 4Claude Opus 4.6
LMArena Vision11921316
GeoBench49%—
VPCT38%—
Furniture Assembly—28.3%
LMArena Document—1507

Multilingual Claude Opus 4.6 leads

Claude Opus 4: 48.8 (#138), Claude Opus 4.6: 57.9 (#6)

Multilingual benchmarks
BenchmarkClaude Opus 4Claude Opus 4.6
LMArena Non-English13621489
LMArena Chinese13861551
LMArena French13721513
LMArena German13911502
LMArena Japanese13311484
LMArena Korean13211464
LMArena Russian13921497
LMArena Spanish13891510

Instruction Following Claude Opus 4.6 leads

Claude Opus 4: 77.1 (#28), Claude Opus 4.6: 79.5 (#4)

Instruction Following benchmarks
BenchmarkClaude Opus 4Claude Opus 4.6
LMArena Instruction Following14061523
IFEval91.8%—

Long Context Claude Opus 4.6 leads

Claude Opus 4: 39.6 (#172), Claude Opus 4.6: 48.1 (#13)

Long Context benchmarks
BenchmarkClaude Opus 4Claude Opus 4.6
LMArena Longer Query14221520
Fiction.LiveBench61.1%—
CL-bench—20.7%
CL-bench Life—17%

Writing & Preference Claude Opus 4.6 leads

Claude Opus 4: 61.2 (#89), Claude Opus 4.6: 73.5 (#10)

Writing & Preference benchmarks
BenchmarkClaude Opus 4Claude Opus 4.6
LMArena Text13771503
LMArena Creative Writing13871505
EQ-Bench Creative Writing15801809
LMArena Multi-Turn13961513
Short-Story Creative Writing83.6%—
WildBench85.2%—
EQ-Bench 4—1223

Frequently asked questions

Is Claude Opus 4 better than Claude Opus 4.6?

Claude Opus 4.6 is the stronger model overall, scoring 58.2 to 43.1 on the Noometry Index.

Which is cheaper, Claude Opus 4 or Claude Opus 4.6?

Claude Opus 4.6 is cheaper. It lists at $5 per million input tokens and $25 per million output tokens; Claude Opus 4 lists at $15 and $75.

Is Claude Opus 4 or Claude Opus 4.6 better for coding?

Claude Opus 4.6 scores higher on coding benchmarks: 57.2 versus 47.2 in the Noometry coding category.

Which has the bigger context window?

Claude Opus 4.6 does, with 1M tokens against 200K.

How many benchmarks do Claude Opus 4 and Claude Opus 4.6 share?

43 benchmarks have published results for both models. Claude Opus 4 has 56 scored results on Noometry and Claude Opus 4.6 has 68.

Related comparisons

Go deeper