Model comparison

Claude Opus 4 vs Gemini 2.5 Pro

Gemini 2.5 Pro is the stronger model overall, scoring 45.0 to 43.1 on the Noometry Index.

Last verified . 55 shared benchmarks.

Claude Opus 4 Anthropic

43.1

Rank #100 Confirmed

Gemini 2.5 Pro Google

45.0

Rank #75 Confirmed

Summary

  • They share 55 benchmarks with published results for both. Claude Opus 4 scores higher in 4 categories and Gemini 2.5 Pro in 6 categories; 10 gaps are clear of the uncertainty.
  • The widest gap is in long context, where Gemini 2.5 Pro leads 59.8 to 39.6.
  • The biggest single-benchmark swing is GeoBench: 49% for Claude Opus 4 and 86% for Gemini 2.5 Pro.
  • Gemini 2.5 Pro is cheaper at $1.25 / $10 per million input/output tokens, against $15 / $75 for Claude Opus 4.
  • Gemini 2.5 Pro accepts more context: 1.05M tokens versus 200K.

Side by side

Claude Opus 4 and Gemini 2.5 Pro specifications
Claude Opus 4Gemini 2.5 Pro
ProviderAnthropicGoogle
Noometry Index43.145.0
Released2025-05-222025-03-25
WeightsProprietaryProprietary
Context window200K1.05M
Max output32K66K
Input $ / M tokens$15$1.25
Output $ / M tokens$75$10
Results tracked5678

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Claude Opus 4 leads

Claude Opus 4: 47.2 (#62), Gemini 2.5 Pro: 42.4 (#101)

Coding benchmarks
BenchmarkClaude Opus 4Gemini 2.5 Pro
SWE-bench Verified70.7%57.6%
SWE-bench Verified (bash only)67.6%53.6%
Aider Polyglot72%83.1%
GSO6.9%3.9%
WeirdML43.7%54%
LMArena Coding14421452
AlgoTune1.331.51
LMArena WebDev—1227
SciCode—42.8%
LiveBench Coding—85.9%
CadEval—64%
ALE-Bench—785.52

Agentic & Tool Use Claude Opus 4 leads

Claude Opus 4: 34.8 (#42), Gemini 2.5 Pro: 29.2 (#88)

Agentic & Tool Use benchmarks
BenchmarkClaude Opus 4Gemini 2.5 Pro
DeepResearch Bench46.8%42.8%
LMArena Search11271142
METR Time Horizons63.9%55.4%
Terminal-Bench—32.6%
GDPval—23.3%
Remote Labor Index—0.8%
TheAgentCompany—30.3%
τ²-bench Banking—13.7%
Cybench38%—
BALROG—43.3%
Vending-Bench 2—573.64

Reasoning Gemini 2.5 Pro leads

Claude Opus 4: 27.3 (#121), Gemini 2.5 Pro: 28.8 (#99)

Reasoning benchmarks
BenchmarkClaude Opus 4Gemini 2.5 Pro
ARC-AGI-28.6%4.9%
SimpleBench58.8%62.4%
Kagi LLM Benchmark74.3%70.3%
ARC-AGI-135.7%41%
CritPt0.3%2%
EnigmaEval5.6%5.6%
LMArena Hard Prompts13991455
DTBench81.6%82.4%
LMCA37.4%34.8%
Epoch Capabilities Index142.67145.32
ForecastBench61.161.3
Chess Puzzles—20%
LiveBench Reasoning—89.8%
LiveBench Data Analysis—79.9%
LiveBench—82.3%

Math Claude Opus 4 leads

Claude Opus 4: 42.0 (#86), Gemini 2.5 Pro: 32.5 (#213)

Math benchmarks
BenchmarkClaude Opus 4Gemini 2.5 Pro
OTIS Mock AIME 2024-202564.4%84.7%
Omni-MATH61.6%41.6%
LMArena Math13901450
MATH Level 585%95.9%
FrontierMath (Feb 2025 set)4.5%14.1%
FrontierMath Tier 4 (v1)4.2%4.2%
FrontierMath (Tiers 1-3)—24.6%
FrontierMath Tier 4—0%
LiveBench Math—90.2%

Knowledge Gemini 2.5 Pro leads

Claude Opus 4: 44.0 (#88), Gemini 2.5 Pro: 56.0 (#46)

Knowledge benchmarks
BenchmarkClaude Opus 4Gemini 2.5 Pro
GPQA Diamond76.3%85.3%
Humanity's Last Exam10.7%21.6%
MMLU-Pro87.5%86.3%
Confabulations15.9%10.6%
Vectara Hallucination Rate12%7%
GPQA (HELM)70.8%74.9%
LMArena Expert13861452

Multimodal Gemini 2.5 Pro leads

Claude Opus 4: 31.5 (#106), Gemini 2.5 Pro: 45.2 (#18)

Multimodal benchmarks
BenchmarkClaude Opus 4Gemini 2.5 Pro
LMArena Vision11921263
GeoBench49%86%
VPCT38%48%
LMArena Document—1421
SpatialViz-Bench—44.7%

Multilingual Gemini 2.5 Pro leads

Claude Opus 4: 48.8 (#138), Gemini 2.5 Pro: 55.3 (#31)

Multilingual benchmarks
BenchmarkClaude Opus 4Gemini 2.5 Pro
LMArena Non-English13621451
LMArena Chinese13861507
LMArena French13721472
LMArena German13911487
LMArena Japanese13311461
LMArena Korean13211434
LMArena Russian13921461
LMArena Spanish13891473

Instruction Following Claude Opus 4 leads

Claude Opus 4: 77.1 (#28), Gemini 2.5 Pro: 75.0 (#75)

Instruction Following benchmarks
BenchmarkClaude Opus 4Gemini 2.5 Pro
IFEval91.8%84%
LMArena Instruction Following14061437
LiveBench Instruction Following—80.6%

Long Context Gemini 2.5 Pro leads

Claude Opus 4: 39.6 (#172), Gemini 2.5 Pro: 59.8 (#5)

Long Context benchmarks
BenchmarkClaude Opus 4Gemini 2.5 Pro
Fiction.LiveBench61.1%91.7%
LMArena Longer Query14221449

Writing & Preference Gemini 2.5 Pro leads

Claude Opus 4: 61.2 (#89), Gemini 2.5 Pro: 63.7 (#62)

Writing & Preference benchmarks
BenchmarkClaude Opus 4Gemini 2.5 Pro
LMArena Text13771458
LMArena Creative Writing13871454
Short-Story Creative Writing83.6%83.8%
EQ-Bench Creative Writing15801421
WildBench85.2%85.7%
LMArena Multi-Turn13961453
LiveBench Language—67.8%

Frequently asked questions

Is Claude Opus 4 better than Gemini 2.5 Pro?

Gemini 2.5 Pro is the stronger model overall, scoring 45.0 to 43.1 on the Noometry Index.

Which is cheaper, Claude Opus 4 or Gemini 2.5 Pro?

Gemini 2.5 Pro is cheaper. It lists at $1.25 per million input tokens and $10 per million output tokens; Claude Opus 4 lists at $15 and $75.

Is Claude Opus 4 or Gemini 2.5 Pro better for coding?

Claude Opus 4 scores higher on coding benchmarks: 47.2 versus 42.4 in the Noometry coding category.

Which has the bigger context window?

Gemini 2.5 Pro does, with 1.05M tokens against 200K.

How many benchmarks do Claude Opus 4 and Gemini 2.5 Pro share?

55 benchmarks have published results for both models. Claude Opus 4 has 56 scored results on Noometry and Gemini 2.5 Pro has 78.

Related comparisons

Go deeper