Model comparison

Claude Opus 4.5 vs Gemini 3 Pro

Gemini 3 Pro is the stronger model overall, scoring 54.8 to 50.5 on the Noometry Index.

Last verified . 59 shared benchmarks.

Claude Opus 4.5 Anthropic

50.5

Rank #47 Confirmed

Gemini 3 Pro Google

54.8

Rank #28 Confirmed

Summary

  • They share 59 benchmarks with published results for both. Claude Opus 4.5 scores higher in 5 categories and Gemini 3 Pro in 5 categories; 10 gaps are clear of the uncertainty.
  • The widest gap is in multimodal, where Gemini 3 Pro leads 57.6 to 31.4.
  • The biggest single-benchmark swing is VPCT: 40% for Claude Opus 4.5 and 91% for Gemini 3 Pro.

Side by side

Claude Opus 4.5 and Gemini 3 Pro specifications
Claude Opus 4.5Gemini 3 Pro
ProviderAnthropicGoogle
Noometry Index50.554.8
Released2025-11-012025-11-18
WeightsProprietaryProprietary
Context window200K—
Max output64K—
Input $ / M tokens$5—
Output $ / M tokens$25—
Results tracked6967

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Claude Opus 4.5 leads

Claude Opus 4.5: 54.8 (#27), Gemini 3 Pro: 51.6 (#39)

Coding benchmarks
BenchmarkClaude Opus 4.5Gemini 3 Pro
SWE-bench Verified76.7%72.9%
SWE-bench Verified (bash only)76.8%74.2%
LMArena WebDev14941440
SWE-bench Multilingual70.7%68.7%
GSO26.5%18.6%
WeirdML63.7%69.9%
LMArena Coding15041481
ALE-Bench1,0251,177
AlgoTune1.771.83

Agentic & Tool Use Claude Opus 4.5 leads

Claude Opus 4.5: 47.3 (#12), Gemini 3 Pro: 40.6 (#23)

Agentic & Tool Use benchmarks
BenchmarkClaude Opus 4.5Gemini 3 Pro
Terminal-Bench63.1%69.4%
Berkeley Function Calling Leaderboard77.5%72.5%
GDPval45.5%40.3%
Remote Labor Index3.8%1.3%
τ²-bench Airline84%80.5%
τ²-bench Banking24.7%18%
τ²-bench Retail79.6%75.9%
τ²-bench Telecom92.3%91%
DeepResearch Bench54.8%46.3%
BALROG43.5%58.1%
LMArena Search11801207
METR Time Horizons75%71%
Vending-Bench 24,9675,478
Cybench82%—
OSWorld66.3%—

Reasoning Gemini 3 Pro leads

Claude Opus 4.5: 42.6 (#51), Gemini 3 Pro: 52.5 (#31)

Reasoning benchmarks
BenchmarkClaude Opus 4.5Gemini 3 Pro
ARC-AGI-237.6%31.1%
SimpleBench62%76.4%
Kagi LLM Benchmark80.2%80.1%
NYT Connections (extended)52.5%94.4%
ARC-AGI-180%75%
Chess Puzzles12%31%
EnigmaEval11.9%18.2%
LMArena Hard Prompts14761480
Epoch Capabilities Index150.09152.92
ForecastBench60.761.2
CritPt—6.9%
EBR-Bench14.3%—
Mystery Game Puzzles22%—
DTBench89.9%—
LMCA44.5%—

Math Gemini 3 Pro leads

Claude Opus 4.5: 38.6 (#132), Gemini 3 Pro: 49.9 (#59)

Math benchmarks
BenchmarkClaude Opus 4.5Gemini 3 Pro
OTIS Mock AIME 2024-202586.1%91.4%
ProofBench36%20%
LMArena Math14631476
FrontierMath (Feb 2025 set)20.7%37.6%
FrontierMath Tier 4 (v1)4.2%18.8%
FrontierMath (Tiers 1-3)34.4%—
FrontierMath Tier 44.9%—
MathArena Final-Answer Competitions—67%
Omni-MATH—55.5%

Knowledge Gemini 3 Pro leads

Claude Opus 4.5: 56.5 (#44), Gemini 3 Pro: 64.4 (#16)

Knowledge benchmarks
BenchmarkClaude Opus 4.5Gemini 3 Pro
GPQA Diamond86%92.6%
Humanity's Last Exam25.2%37.5%
Vectara Hallucination Rate10.9%13.6%
LMArena Expert14871475
SimpleQA Verified45.7%—
MMLU-Pro—90.3%
GPQA (HELM)—80.3%

Multimodal Gemini 3 Pro leads

Claude Opus 4.5: 31.4 (#107), Gemini 3 Pro: 57.6 (#2)

Multimodal benchmarks
BenchmarkClaude Opus 4.5Gemini 3 Pro
GeoBench75%84%
VPCT40%91%
LMArena Document14621434
LMArena Vision—1305
Furniture Assembly28.3%—

Multilingual Gemini 3 Pro leads

Claude Opus 4.5: 54.3 (#47), Gemini 3 Pro: 56.9 (#16)

Multilingual benchmarks
BenchmarkClaude Opus 4.5Gemini 3 Pro
LMArena Non-English14381474
LMArena Chinese14701523
LMArena French14711492
LMArena German14491515
LMArena Japanese14161510
LMArena Korean14241448
LMArena Russian14471493
LMArena Spanish14581470

Instruction Following Claude Opus 4.5 leads

Claude Opus 4.5: 77.5 (#19), Gemini 3 Pro: 76.3 (#45)

Instruction Following benchmarks
BenchmarkClaude Opus 4.5Gemini 3 Pro
LMArena Instruction Following14781458
IFEval—87.7%

Long Context Claude Opus 4.5 leads

Claude Opus 4.5: 46.5 (#22), Gemini 3 Pro: 44.0 (#79)

Long Context benchmarks
BenchmarkClaude Opus 4.5Gemini 3 Pro
CL-bench21.1%15.8%
LMArena Longer Query14801471

Writing & Preference Claude Opus 4.5 leads

Claude Opus 4.5: 68.1 (#28), Gemini 3 Pro: 66.4 (#35)

Writing & Preference benchmarks
BenchmarkClaude Opus 4.5Gemini 3 Pro
LMArena Text14511479
LMArena Creative Writing14451482
EQ-Bench Creative Writing16871525
LMArena Multi-Turn14661484
WildBench—85.9%

Frequently asked questions

Is Claude Opus 4.5 better than Gemini 3 Pro?

Gemini 3 Pro is the stronger model overall, scoring 54.8 to 50.5 on the Noometry Index.

Is Claude Opus 4.5 or Gemini 3 Pro better for coding?

Claude Opus 4.5 scores higher on coding benchmarks: 54.8 versus 51.6 in the Noometry coding category.

How many benchmarks do Claude Opus 4.5 and Gemini 3 Pro share?

59 benchmarks have published results for both models. Claude Opus 4.5 has 69 scored results on Noometry and Gemini 3 Pro has 67.

Related comparisons

Go deeper