Model comparison

Gemini 3 Pro vs o1

Gemini 3 Pro is the stronger model overall, scoring 54.8 to 40.9 on the Noometry Index.

Last verified . 31 shared benchmarks.

Gemini 3 Pro Google

54.8

Rank #28 Confirmed

o1 OpenAI

40.9

Rank #143 Confirmed

Summary

  • They share 31 benchmarks with published results for both. Gemini 3 Pro scores higher in 9 categories and o1 in 1 category; 10 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Gemini 3 Pro leads 52.5 to 27.9.
  • The biggest single-benchmark swing is VPCT: 91% for Gemini 3 Pro and 37% for o1.

Side by side

Gemini 3 Pro and o1 specifications
Gemini 3 Proo1
ProviderGoogleOpenAI
Noometry Index54.840.9
Released2025-11-182024-09-12
WeightsProprietaryProprietary
Context window—200K
Max output—100K
Input $ / M tokens—$15
Output $ / M tokens—$60
Results tracked6752

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Gemini 3 Pro leads

Gemini 3 Pro: 51.6 (#39), o1: 46.1 (#70)

Coding benchmarks
BenchmarkGemini 3 Proo1
WeirdML69.9%47.6%
LMArena Coding14811367
SWE-bench Verified72.9%—
SWE-bench Verified (bash only)74.2%—
Aider Polyglot—61.7%
LMArena WebDev1440—
SWE-bench Multilingual68.7%—
GSO18.6%—
LiveBench Coding—69.7%
CadEval—56%
ALE-Bench1,177—
AlgoTune1.83—
HumanEval+—89%
MBPP+—80.2%

Agentic & Tool Use Gemini 3 Pro leads

Gemini 3 Pro: 40.6 (#23), o1: 24.6 (#117)

Agentic & Tool Use benchmarks
BenchmarkGemini 3 Proo1
METR Time Horizons71%51.1%
Terminal-Bench69.4%—
Berkeley Function Calling Leaderboard72.5%—
GDPval40.3%—
Remote Labor Index1.3%—
τ²-bench Airline80.5%—
τ²-bench Banking18%—
τ²-bench Retail75.9%—
τ²-bench Telecom91%—
Cybench—10%
DeepResearch Bench46.3%—
BALROG58.1%—
LMArena Search1207—
Vending-Bench 25,478—

Reasoning Gemini 3 Pro leads

Gemini 3 Pro: 52.5 (#31), o1: 27.9 (#111)

Reasoning benchmarks
BenchmarkGemini 3 Proo1
SimpleBench76.4%41.7%
ARC-AGI-175%30.7%
Chess Puzzles31%15%
EnigmaEval18.2%5.7%
LMArena Hard Prompts14801371
Epoch Capabilities Index152.92141.91
ARC-AGI-231.1%—
Kagi LLM Benchmark80.1%—
NYT Connections (extended)94.4%—
CritPt6.9%—
LiveBench Reasoning—91.6%
DTBench—74.7%
LiveBench Data Analysis—65.5%
LMCA—22.3%
ForecastBench61.2—
LiveBench—75.7%

Math Gemini 3 Pro leads

Gemini 3 Pro: 49.9 (#59), o1: 36.1 (#175)

Knowledge Gemini 3 Pro leads

Gemini 3 Pro: 64.4 (#16), o1: 41.5 (#110)

Knowledge benchmarks
BenchmarkGemini 3 Proo1
GPQA Diamond92.6%76.8%
Humanity's Last Exam37.5%8%
LMArena Expert14751361
SimpleQA Verified—41.1%
MMLU-Pro90.3%—
Confabulations—11.7%
Vectara Hallucination Rate13.6%—
GPQA (HELM)80.3%—

Multimodal Gemini 3 Pro leads

Gemini 3 Pro: 57.6 (#2), o1: 34.2 (#93)

Multimodal benchmarks
BenchmarkGemini 3 Proo1
LMArena Vision13051168
GeoBench84%80%
VPCT91%37%
LMArena Document1434—
SpatialViz-Bench—41.4%

Multilingual Gemini 3 Pro leads

Gemini 3 Pro: 56.9 (#16), o1: 48.6 (#142)

Multilingual benchmarks
BenchmarkGemini 3 Proo1
LMArena Non-English14741358
LMArena Chinese15231394
LMArena French14921344
LMArena German15151337
LMArena Japanese15101346
LMArena Korean14481396
LMArena Russian14931356
LMArena Spanish14701345

Instruction Following Gemini 3 Pro leads

Gemini 3 Pro: 76.3 (#45), o1: 74.8 (#86)

Instruction Following benchmarks
BenchmarkGemini 3 Proo1
LMArena Instruction Following14581367
LiveBench Instruction Following—81.5%
IFEval87.7%—

Long Context o1 leads

Gemini 3 Pro: 44.0 (#79), o1: 50.3 (#9)

Long Context benchmarks
BenchmarkGemini 3 Proo1
LMArena Longer Query14711378
Fiction.LiveBench—83.3%
CL-bench15.8%—

Writing & Preference Gemini 3 Pro leads

Gemini 3 Pro: 66.4 (#35), o1: 55.6 (#144)

Writing & Preference benchmarks
BenchmarkGemini 3 Proo1
LMArena Text14791366
LMArena Creative Writing14821348
LMArena Multi-Turn14841369
Short-Story Creative Writing—70.2%
EQ-Bench Creative Writing1525—
WildBench85.9%—
LiveBench Language—65.4%

Frequently asked questions

Is Gemini 3 Pro better than o1?

Gemini 3 Pro is the stronger model overall, scoring 54.8 to 40.9 on the Noometry Index.

Is Gemini 3 Pro or o1 better for coding?

Gemini 3 Pro scores higher on coding benchmarks: 51.6 versus 46.1 in the Noometry coding category.

How many benchmarks do Gemini 3 Pro and o1 share?

31 benchmarks have published results for both models. Gemini 3 Pro has 67 scored results on Noometry and o1 has 52.

Related comparisons

Go deeper