Model comparison

Claude Opus 4 vs Claude Opus 4.1

Claude Opus 4 is the stronger model overall, scoring 43.1 to 41.0 on the Noometry Index.

Last verified . 39 shared benchmarks.

Claude Opus 4 Anthropic

43.1

Rank #100 Confirmed

Claude Opus 4.1 Anthropic

41.0

Rank #142 Confirmed

Summary

  • They share 39 benchmarks with published results for both. Claude Opus 4 scores higher in 5 categories and Claude Opus 4.1 in 5 categories; 9 gaps are clear of the uncertainty.
  • The widest gap is in math, where Claude Opus 4 leads 42.0 to 22.3.
  • Both cost about the same: $15 input and $75 output per million tokens.

Side by side

Claude Opus 4 and Claude Opus 4.1 specifications
Claude Opus 4Claude Opus 4.1
ProviderAnthropicAnthropic
Noometry Index43.141.0
Released2025-05-222025-08-05
WeightsProprietaryProprietary
Context window200K200K
Max output32K32K
Input $ / M tokens$15$15
Output $ / M tokens$75$75
Results tracked5648

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Claude Opus 4 leads

Claude Opus 4: 47.2 (#62), Claude Opus 4.1: 44.4 (#73)

Coding benchmarks
BenchmarkClaude Opus 4Claude Opus 4.1
SWE-bench Verified70.7%73.3%
WeirdML43.7%45.9%
LMArena Coding14421479
AlgoTune1.331.34
SWE-bench Verified (bash only)67.6%—
Aider Polyglot72%—
LMArena WebDev—1390
GSO6.9%—
ALE-Bench—674.77

Agentic & Tool Use Too close to call

Claude Opus 4: 34.8 (#42), Claude Opus 4.1: 35.0 (#41)

Agentic & Tool Use benchmarks
BenchmarkClaude Opus 4Claude Opus 4.1
Cybench38%42%
DeepResearch Bench46.8%48.3%
LMArena Search11271148
METR Time Horizons63.9%66.8%
Terminal-Bench—38%
GDPval—43.6%

Reasoning Claude Opus 4.1 leads

Claude Opus 4: 27.3 (#121), Claude Opus 4.1: 32.2 (#76)

Reasoning benchmarks
BenchmarkClaude Opus 4Claude Opus 4.1
SimpleBench58.8%60%
EnigmaEval5.6%7.2%
LMArena Hard Prompts13991443
DTBench81.6%80%
LMCA37.4%37.1%
Epoch Capabilities Index142.67144.12
ForecastBench61.162
ARC-AGI-28.6%—
Kagi LLM Benchmark74.3%—
ARC-AGI-135.7%—
CritPt0.3%—
Chess Puzzles—7%
EBR-Bench—7.9%
Mystery Game Puzzles—21%

Math Claude Opus 4 leads

Claude Opus 4: 42.0 (#86), Claude Opus 4.1: 22.3 (#277)

Math benchmarks
BenchmarkClaude Opus 4Claude Opus 4.1
OTIS Mock AIME 2024-202564.4%68.9%
LMArena Math13901431
FrontierMath (Feb 2025 set)4.5%7.2%
FrontierMath Tier 4 (v1)4.2%4.2%
FrontierMath (Tiers 1-3)—12.6%
FrontierMath Tier 4—2.4%
Omni-MATH61.6%—
MATH Level 585%—

Knowledge Claude Opus 4 leads

Claude Opus 4: 44.0 (#88), Claude Opus 4.1: 42.0 (#101)

Knowledge benchmarks
BenchmarkClaude Opus 4Claude Opus 4.1
GPQA Diamond76.3%77.3%
Humanity's Last Exam10.7%11.5%
Confabulations15.9%17.1%
Vectara Hallucination Rate12%11.8%
LMArena Expert13861439
MMLU-Pro87.5%—
GPQA (HELM)70.8%—

Multimodal Claude Opus 4 leads

Claude Opus 4: 31.5 (#106), Claude Opus 4.1: 26.8 (#119)

Multimodal benchmarks
BenchmarkClaude Opus 4Claude Opus 4.1
VPCT38%35%
LMArena Vision1192—
GeoBench49%—

Multilingual Claude Opus 4.1 leads

Claude Opus 4: 48.8 (#138), Claude Opus 4.1: 52.0 (#95)

Multilingual benchmarks
BenchmarkClaude Opus 4Claude Opus 4.1
LMArena Non-English13621405
LMArena Chinese13861427
LMArena French13721431
LMArena German13911413
LMArena Japanese13311378
LMArena Korean13211380
LMArena Russian13921422
LMArena Spanish13891448

Instruction Following Claude Opus 4 leads

Claude Opus 4: 77.1 (#28), Claude Opus 4.1: 75.6 (#58)

Instruction Following benchmarks
BenchmarkClaude Opus 4Claude Opus 4.1
LMArena Instruction Following14061435
IFEval91.8%—

Long Context Claude Opus 4.1 leads

Claude Opus 4: 39.6 (#172), Claude Opus 4.1: 44.5 (#63)

Long Context benchmarks
BenchmarkClaude Opus 4Claude Opus 4.1
LMArena Longer Query14221455
Fiction.LiveBench61.1%—

Writing & Preference Claude Opus 4.1 leads

Claude Opus 4: 61.2 (#89), Claude Opus 4.1: 62.4 (#74)

Writing & Preference benchmarks
BenchmarkClaude Opus 4Claude Opus 4.1
LMArena Text13771419
LMArena Creative Writing13871412
Short-Story Creative Writing83.6%84.7%
LMArena Multi-Turn13961444
EQ-Bench Creative Writing1580—
WildBench85.2%—

Frequently asked questions

Is Claude Opus 4 better than Claude Opus 4.1?

Claude Opus 4 is the stronger model overall, scoring 43.1 to 41.0 on the Noometry Index.

Which is cheaper, Claude Opus 4 or Claude Opus 4.1?

Claude Opus 4.1 is cheaper. It lists at $15 per million input tokens and $75 per million output tokens; Claude Opus 4 lists at $15 and $75.

Is Claude Opus 4 or Claude Opus 4.1 better for coding?

Claude Opus 4 scores higher on coding benchmarks: 47.2 versus 44.4 in the Noometry coding category.

Which has the bigger context window?

Both accept 200K tokens.

How many benchmarks do Claude Opus 4 and Claude Opus 4.1 share?

39 benchmarks have published results for both models. Claude Opus 4 has 56 scored results on Noometry and Claude Opus 4.1 has 48.

Related comparisons

Go deeper