Model comparison

Gemini 3 Pro vs o4-mini

Gemini 3 Pro is the stronger model overall, scoring 54.8 to 41.6 on the Noometry Index.

Last verified . 48 shared benchmarks.

Gemini 3 Pro Google

54.8

Rank #28 Confirmed

o4-mini OpenAI

41.6

Rank #132 Confirmed

Summary

  • They share 48 benchmarks with published results for both. Gemini 3 Pro scores higher in 9 categories and o4-mini in 1 category; 10 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Gemini 3 Pro leads 52.5 to 24.6.
  • The biggest single-benchmark swing is SimpleBench: 76.4% for Gemini 3 Pro and 38.7% for o4-mini.

Side by side

Gemini 3 Pro and o4-mini specifications
Gemini 3 Proo4-mini
ProviderGoogleOpenAI
Noometry Index54.841.6
Released2025-11-182025-04-16
WeightsProprietaryProprietary
Context window—200K
Max output—100K
Input $ / M tokens—$1.10
Output $ / M tokens—$4.40
Results tracked6760

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Gemini 3 Pro leads

Gemini 3 Pro: 51.6 (#39), o4-mini: 40.9 (#127)

Coding benchmarks
BenchmarkGemini 3 Proo4-mini
SWE-bench Verified (bash only)74.2%45%
GSO18.6%3.6%
WeirdML69.9%52.6%
LMArena Coding14811368
ALE-Bench1,177826.17
AlgoTune1.831.72
SWE-bench Verified72.9%—
Aider Polyglot—72%
LMArena WebDev1440—
SWE-bench Multilingual68.7%—
CadEval—62%

Agentic & Tool Use Gemini 3 Pro leads

Gemini 3 Pro: 40.6 (#23), o4-mini: 32.6 (#61)

Agentic & Tool Use benchmarks
BenchmarkGemini 3 Proo4-mini
Berkeley Function Calling Leaderboard72.5%53.2%
GDPval40.3%25.3%
METR Time Horizons71%63.9%
Terminal-Bench69.4%—
Remote Labor Index1.3%—
τ²-bench Airline80.5%—
τ²-bench Banking18%—
τ²-bench Retail75.9%—
τ²-bench Telecom91%—
DeepResearch Bench46.3%—
BALROG58.1%—
LMArena Search1207—
Vending-Bench 25,478—

Reasoning Gemini 3 Pro leads

Gemini 3 Pro: 52.5 (#31), o4-mini: 24.6 (#162)

Reasoning benchmarks
BenchmarkGemini 3 Proo4-mini
ARC-AGI-231.1%6.1%
SimpleBench76.4%38.7%
Kagi LLM Benchmark80.1%67.6%
ARC-AGI-175%58.7%
CritPt6.9%0.6%
Chess Puzzles31%26%
EnigmaEval18.2%9.2%
LMArena Hard Prompts14801351
Epoch Capabilities Index152.92145.64
ForecastBench61.261.8
NYT Connections (extended)94.4%—
Mystery Game Puzzles—5%
DTBench—77.6%
LMCA—26.5%

Math Gemini 3 Pro leads

Gemini 3 Pro: 49.9 (#59), o4-mini: 40.8 (#89)

Knowledge Gemini 3 Pro leads

Gemini 3 Pro: 64.4 (#16), o4-mini: 43.6 (#91)

Knowledge benchmarks
BenchmarkGemini 3 Proo4-mini
GPQA Diamond92.6%79.6%
Humanity's Last Exam37.5%18.1%
MMLU-Pro90.3%82%
Vectara Hallucination Rate13.6%18.6%
GPQA (HELM)80.3%73.5%
LMArena Expert14751343
SimpleQA Verified—19.6%
Confabulations—15.8%

Multimodal Gemini 3 Pro leads

Gemini 3 Pro: 57.6 (#2), o4-mini: 40.2 (#49)

Multimodal benchmarks
BenchmarkGemini 3 Proo4-mini
LMArena Vision13051194
GeoBench84%64%
VPCT91%57.5%
LMArena Document1434—

Multilingual Gemini 3 Pro leads

Gemini 3 Pro: 56.9 (#16), o4-mini: 47.0 (#154)

Multilingual benchmarks
BenchmarkGemini 3 Proo4-mini
LMArena Non-English14741337
LMArena Chinese15231354
LMArena French14921364
LMArena German15151336
LMArena Japanese15101308
LMArena Korean14481312
LMArena Russian14931334
LMArena Spanish14701347

Instruction Following Gemini 3 Pro leads

Gemini 3 Pro: 76.3 (#45), o4-mini: 75.2 (#68)

Instruction Following benchmarks
BenchmarkGemini 3 Proo4-mini
IFEval87.7%92.8%
LMArena Instruction Following14581321

Long Context o4-mini leads

Gemini 3 Pro: 44.0 (#79), o4-mini: 45.5 (#33)

Long Context benchmarks
BenchmarkGemini 3 Proo4-mini
LMArena Longer Query14711315
Fiction.LiveBench—77.8%
CL-bench15.8%—

Writing & Preference Gemini 3 Pro leads

Gemini 3 Pro: 66.4 (#35), o4-mini: 54.0 (#152)

Writing & Preference benchmarks
BenchmarkGemini 3 Proo4-mini
LMArena Text14791353
LMArena Creative Writing14821294
WildBench85.9%85.4%
LMArena Multi-Turn14841350
Short-Story Creative Writing—75%
EQ-Bench Creative Writing1525—

Frequently asked questions

Is Gemini 3 Pro better than o4-mini?

Gemini 3 Pro is the stronger model overall, scoring 54.8 to 41.6 on the Noometry Index.

Is Gemini 3 Pro or o4-mini better for coding?

Gemini 3 Pro scores higher on coding benchmarks: 51.6 versus 40.9 in the Noometry coding category.

How many benchmarks do Gemini 3 Pro and o4-mini share?

48 benchmarks have published results for both models. Gemini 3 Pro has 67 scored results on Noometry and o4-mini has 60.

Related comparisons

Go deeper