Model comparison

Claude Opus 4.8 vs Gemini 3 Pro

Claude Opus 4.8 is the stronger model overall, scoring 60.7 to 54.8 on the Noometry Index.

Last verified . 45 shared benchmarks.

Claude Opus 4.8 Anthropic

60.7

Rank #13 Confirmed

Gemini 3 Pro Google

54.8

Rank #28 Confirmed

Summary

  • They share 45 benchmarks with published results for both. Claude Opus 4.8 scores higher in 7 categories and Gemini 3 Pro in 3 categories; 10 gaps are clear of the uncertainty.
  • The widest gap is in math, where Claude Opus 4.8 leads 78.4 to 49.9.
  • The biggest single-benchmark swing is ProofBench: 69% for Claude Opus 4.8 and 20% for Gemini 3 Pro.

Side by side

Claude Opus 4.8 and Gemini 3 Pro specifications
Claude Opus 4.8Gemini 3 Pro
ProviderAnthropicGoogle
Noometry Index60.754.8
Released2026-05-282025-11-18
WeightsProprietaryProprietary
Context window1M—
Max output128K—
Input $ / M tokens$5—
Output $ / M tokens$25—
Results tracked6567

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Claude Opus 4.8 leads

Claude Opus 4.8: 59.9 (#12), Gemini 3 Pro: 51.6 (#39)

Coding benchmarks
BenchmarkClaude Opus 4.8Gemini 3 Pro
LMArena WebDev15561440
GSO47.1%18.6%
WeirdML82.9%69.9%
LMArena Coding14901481
ALE-Bench1,5641,177
SWE-bench Verified—72.9%
DeepSWE59%—
FrontierCode46.5%—
SWE-bench Verified (bash only)—74.2%
SWE-bench Multilingual—68.7%
SciCode53.5%—
AlgoTune—1.83

Agentic & Tool Use Claude Opus 4.8 leads

Claude Opus 4.8: 47.6 (#11), Gemini 3 Pro: 40.6 (#23)

Agentic & Tool Use benchmarks
BenchmarkClaude Opus 4.8Gemini 3 Pro
Remote Labor Index8.3%1.3%
τ²-bench Banking39.7%18%
DeepResearch Bench50.2%46.3%
LMArena Search12041207
Vending-Bench 25,7875,478
Terminal-Bench—69.4%
APEX-Agents48.9%—
Berkeley Function Calling Leaderboard—72.5%
OSWorld 2.020.6%—
GDPval—40.3%
τ²-bench Airline—80.5%
τ²-bench Retail—75.9%
τ²-bench Telecom—91%
PostTrainBench33.8%—
BALROG—58.1%
GBAEval70.9%—
GDP.pdf24%—
METR Time Horizons—71%

Reasoning Claude Opus 4.8 leads

Claude Opus 4.8: 64.7 (#16), Gemini 3 Pro: 52.5 (#31)

Reasoning benchmarks
BenchmarkClaude Opus 4.8Gemini 3 Pro
ARC-AGI-272.1%31.1%
SimpleBench64.8%76.4%
Kagi LLM Benchmark88.8%80.1%
NYT Connections (extended)91.1%94.4%
ARC-AGI-192.5%75%
CritPt20.9%6.9%
Chess Puzzles34%31%
EnigmaEval23.5%18.2%
LMArena Hard Prompts14821480
Epoch Capabilities Index158.21152.92
ForecastBench59.961.2
EBR-Bench28.6%—
Mystery Game Puzzles36%—
DTBench94.9%—
LMCA57.5%—
Surface Evolver Bench87.5%—
Bench to the Future 30.14—

Math Claude Opus 4.8 leads

Claude Opus 4.8: 78.4 (#13), Gemini 3 Pro: 49.9 (#59)

Math benchmarks
BenchmarkClaude Opus 4.8Gemini 3 Pro
MathArena Final-Answer Competitions91.8%67%
OTIS Mock AIME 2024-202598.3%91.4%
ProofBench69%20%
LMArena Math14871476
FrontierMath (Feb 2025 set)47.2%37.6%
FrontierMath Tier 4 (v1)31.3%18.8%
FrontierMath (Tiers 1-3)80%—
FrontierMath Tier 456.1%—
Omni-MATH—55.5%

Knowledge Gemini 3 Pro leads

Claude Opus 4.8: 61.3 (#29), Gemini 3 Pro: 64.4 (#16)

Knowledge benchmarks
BenchmarkClaude Opus 4.8Gemini 3 Pro
GPQA Diamond91%92.6%
LMArena Expert15021475
Humanity's Last Exam—37.5%
SimpleQA Verified53%—
MMLU-Pro—90.3%
Vectara Hallucination Rate—13.6%
GPQA (HELM)—80.3%

Multimodal Gemini 3 Pro leads

Claude Opus 4.8: 42.9 (#26), Gemini 3 Pro: 57.6 (#2)

Multimodal benchmarks
BenchmarkClaude Opus 4.8Gemini 3 Pro
LMArena Vision12941305
LMArena Document14751434
GeoBench—84%
VPCT—91%
Blueprint-Bench 214.5%—
Furniture Assembly42.5%—

Multilingual Gemini 3 Pro leads

Claude Opus 4.8: 55.2 (#33), Gemini 3 Pro: 56.9 (#16)

Multilingual benchmarks
BenchmarkClaude Opus 4.8Gemini 3 Pro
LMArena Non-English14501474
LMArena Chinese15071523
LMArena French14811492
LMArena German14721515
LMArena Japanese14401510
LMArena Korean14321448
LMArena Russian14741493
LMArena Spanish14661470

Instruction Following Claude Opus 4.8 leads

Claude Opus 4.8: 77.4 (#24), Gemini 3 Pro: 76.3 (#45)

Instruction Following benchmarks
BenchmarkClaude Opus 4.8Gemini 3 Pro
LMArena Instruction Following14761458
IFEval—87.7%

Long Context Claude Opus 4.8 leads

Claude Opus 4.8: 45.4 (#35), Gemini 3 Pro: 44.0 (#79)

Long Context benchmarks
BenchmarkClaude Opus 4.8Gemini 3 Pro
LMArena Longer Query14831471
CL-bench—15.8%

Writing & Preference Claude Opus 4.8 leads

Claude Opus 4.8: 72.0 (#16), Gemini 3 Pro: 66.4 (#35)

Writing & Preference benchmarks
BenchmarkClaude Opus 4.8Gemini 3 Pro
LMArena Text14611479
LMArena Creative Writing14541482
EQ-Bench Creative Writing18401525
LMArena Multi-Turn14761484
WildBench—85.9%
EQ-Bench 41281—

Frequently asked questions

Is Claude Opus 4.8 better than Gemini 3 Pro?

Claude Opus 4.8 is the stronger model overall, scoring 60.7 to 54.8 on the Noometry Index.

Is Claude Opus 4.8 or Gemini 3 Pro better for coding?

Claude Opus 4.8 scores higher on coding benchmarks: 59.9 versus 51.6 in the Noometry coding category.

How many benchmarks do Claude Opus 4.8 and Gemini 3 Pro share?

45 benchmarks have published results for both models. Claude Opus 4.8 has 65 scored results on Noometry and Gemini 3 Pro has 67.

Related comparisons

Go deeper