Model comparison

Claude Opus 5.5 vs Gemini 3 Pro

Claude Opus 5.5 is the stronger model overall, scoring 68.6 to 54.8 on the Noometry Index.

Last verified . 27 shared benchmarks.

Claude Opus 5.5 Anthropic

68.6

Rank #3 Confirmed

Gemini 3 Pro Google

54.8

Rank #28 Confirmed

Summary

  • They share 27 benchmarks with published results for both. Claude Opus 5.5 scores higher in 10 categories and Gemini 3 Pro in 0 categories; 9 gaps are clear of the uncertainty.
  • The widest gap is in math, where Claude Opus 5.5 leads 91.8 to 49.9.
  • The biggest single-benchmark swing is ProofBench: 100% for Claude Opus 5.5 and 20% for Gemini 3 Pro.

Side by side

Claude Opus 5.5 and Gemini 3 Pro specifications
Claude Opus 5.5Gemini 3 Pro
ProviderAnthropicGoogle
Noometry Index68.654.8
Released2026-09-222025-11-18
WeightsProprietaryProprietary
Context window1M—
Max output128K—
Input $ / M tokens$4—
Output $ / M tokens$20—
Results tracked4467

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Claude Opus 5.5 leads

Claude Opus 5.5: 71.9 (#3), Gemini 3 Pro: 51.6 (#39)

Coding benchmarks
BenchmarkClaude Opus 5.5Gemini 3 Pro
LMArena WebDev18131440
LMArena Coding15471481
ALE-Bench2,1471,177
SWE-bench Verified—72.9%
FrontierCode54.6%—
SWE-bench Verified (bash only)—74.2%
CursorBench57.8%—
SWE-bench Multilingual—68.7%
FrontierSWE62.3%—
SciCode66.9%—
GSO—18.6%
WeirdML—69.9%
MirrorCode77.4%—
AlgoTune—1.83

Agentic & Tool Use Claude Opus 5.5 leads

Claude Opus 5.5: 45.3 (#15), Gemini 3 Pro: 40.6 (#23)

Agentic & Tool Use benchmarks
BenchmarkClaude Opus 5.5Gemini 3 Pro
Vending-Bench 29,2355,478
Terminal-Bench—69.4%
APEX-Agents73.5%—
Berkeley Function Calling Leaderboard—72.5%
GDPval—40.3%
Remote Labor Index—1.3%
τ²-bench Airline—80.5%
τ²-bench Banking—18%
τ²-bench Retail—75.9%
τ²-bench Telecom—91%
DeepResearch Bench—46.3%
BALROG—58.1%
GDP.pdf30.6%—
LMArena Search—1207
METR Time Horizons—71%

Reasoning Claude Opus 5.5 leads

Claude Opus 5.5: 80.2 (#3), Gemini 3 Pro: 52.5 (#31)

Reasoning benchmarks
BenchmarkClaude Opus 5.5Gemini 3 Pro
ARC-AGI-293.3%31.1%
NYT Connections (extended)88.5%94.4%
ARC-AGI-198.5%75%
CritPt31.7%6.9%
LMArena Hard Prompts15351480
Epoch Capabilities Index167.33152.92
SimpleBench—76.4%
Kagi LLM Benchmark—80.1%
Chess Puzzles—31%
EnigmaEval—18.2%
EBR-Bench71.4%—
Mystery Game Puzzles71%—
DTBench98.9%—
LMCA68.2%—
ForecastBench—61.2

Math Claude Opus 5.5 leads

Claude Opus 5.5: 91.8 (#3), Gemini 3 Pro: 49.9 (#59)

Knowledge Claude Opus 5.5 leads

Claude Opus 5.5: 66.4 (#10), Gemini 3 Pro: 64.4 (#16)

Knowledge benchmarks
BenchmarkClaude Opus 5.5Gemini 3 Pro
GPQA Diamond90.6%92.6%
LMArena Expert15471475
Humanity's Last Exam—37.5%
SimpleQA Verified72.2%—
MMLU-Pro—90.3%
Vectara Hallucination Rate—13.6%
GPQA (HELM)—80.3%

Multimodal Too close to call

Claude Opus 5.5: 57.8 (#1), Gemini 3 Pro: 57.6 (#2)

Multimodal benchmarks
BenchmarkClaude Opus 5.5Gemini 3 Pro
LMArena Vision13211305
GeoBench—84%
VPCT—91%
Blueprint-Bench 251.2%—
Furniture Assembly83.3%—
LMArena Document—1434

Multilingual Claude Opus 5.5 leads

Claude Opus 5.5: 59.1 (#2), Gemini 3 Pro: 56.9 (#16)

Multilingual benchmarks
BenchmarkClaude Opus 5.5Gemini 3 Pro
LMArena Non-English15071474
LMArena Chinese15881523
LMArena French15141492
LMArena Russian15201493
LMArena Spanish15071470
LMArena German—1515
LMArena Japanese—1510
LMArena Korean—1448

Instruction Following Claude Opus 5.5 leads

Claude Opus 5.5: 80.0 (#3), Gemini 3 Pro: 76.3 (#45)

Instruction Following benchmarks
BenchmarkClaude Opus 5.5Gemini 3 Pro
LMArena Instruction Following15371458
IFEval—87.7%

Long Context Claude Opus 5.5 leads

Claude Opus 5.5: 47.1 (#19), Gemini 3 Pro: 44.0 (#79)

Long Context benchmarks
BenchmarkClaude Opus 5.5Gemini 3 Pro
LMArena Longer Query15321471
CL-bench—15.8%

Writing & Preference Claude Opus 5.5 leads

Claude Opus 5.5: 78.2 (#3), Gemini 3 Pro: 66.4 (#35)

Writing & Preference benchmarks
BenchmarkClaude Opus 5.5Gemini 3 Pro
LMArena Text15151479
LMArena Creative Writing15331482
EQ-Bench Creative Writing20501525
LMArena Multi-Turn14991484
WildBench—85.9%

Frequently asked questions

Is Claude Opus 5.5 better than Gemini 3 Pro?

Claude Opus 5.5 is the stronger model overall, scoring 68.6 to 54.8 on the Noometry Index.

Is Claude Opus 5.5 or Gemini 3 Pro better for coding?

Claude Opus 5.5 scores higher on coding benchmarks: 71.9 versus 51.6 in the Noometry coding category.

How many benchmarks do Claude Opus 5.5 and Gemini 3 Pro share?

27 benchmarks have published results for both models. Claude Opus 5.5 has 44 scored results on Noometry and Gemini 3 Pro has 67.

Related comparisons

Go deeper