Model comparison

Gemini 2.5 Pro vs GPT-4.5

Gemini 2.5 Pro is the stronger model overall, scoring 45.0 to 37.2 on the Noometry Index.

Last verified . 41 shared benchmarks.

Gemini 2.5 Pro Google

45.0

Rank #75 Confirmed

GPT-4.5 OpenAI

37.2

Rank #208 Confirmed

Summary

  • They share 41 benchmarks with published results for both. Gemini 2.5 Pro scores higher in 9 categories and GPT-4.5 in 1 category; 8 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Gemini 2.5 Pro leads 56.0 to 32.5.
  • The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 84.7% for Gemini 2.5 Pro and 37.8% for GPT-4.5.

Side by side

Gemini 2.5 Pro and GPT-4.5 specifications
Gemini 2.5 ProGPT-4.5
ProviderGoogleOpenAI
Noometry Index45.037.2
Released2025-03-252025-02-27
WeightsProprietaryProprietary
Context window1.05M—
Max output66K—
Input $ / M tokens$1.25—
Output $ / M tokens$10—
Results tracked7842

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

Gemini 2.5 Pro: 42.4 (#101), GPT-4.5: 42.2 (#109)

Coding benchmarks
BenchmarkGemini 2.5 ProGPT-4.5
Aider Polyglot83.1%44.9%
WeirdML54%39.4%
LiveBench Coding85.9%75.2%
LMArena Coding14521396
SWE-bench Verified57.6%—
SWE-bench Verified (bash only)53.6%—
LMArena WebDev1227—
SciCode42.8%—
GSO3.9%—
CadEval64%—
ALE-Bench785.52—
AlgoTune1.51—

Agentic & Tool Use Gemini 2.5 Pro leads

Gemini 2.5 Pro: 29.2 (#88), GPT-4.5: 27.9 (#97)

Agentic & Tool Use benchmarks
BenchmarkGemini 2.5 ProGPT-4.5
Terminal-Bench32.6%—
GDPval23.3%—
Remote Labor Index0.8%—
TheAgentCompany30.3%—
τ²-bench Banking13.7%—
Cybench—17.5%
DeepResearch Bench42.8%—
BALROG43.3%—
LMArena Search1142—
METR Time Horizons55.4%—
Vending-Bench 2573.64—

Reasoning Gemini 2.5 Pro leads

Gemini 2.5 Pro: 28.8 (#99), GPT-4.5: 13.9 (#330)

Reasoning benchmarks
BenchmarkGemini 2.5 ProGPT-4.5
ARC-AGI-24.9%0.8%
SimpleBench62.4%34.5%
ARC-AGI-141%10.3%
EnigmaEval5.6%3.2%
LiveBench Reasoning89.8%71.1%
LMArena Hard Prompts14551403
LiveBench Data Analysis79.9%64.3%
Epoch Capabilities Index145.32136.74
ForecastBench61.361.7
LiveBench82.3%69%
Kagi LLM Benchmark70.3%—
CritPt2%—
Chess Puzzles20%—
DTBench82.4%—
LMCA34.8%—

Math Too close to call

Gemini 2.5 Pro: 32.5 (#213), GPT-4.5: 32.6 (#211)

Math benchmarks
BenchmarkGemini 2.5 ProGPT-4.5
OTIS Mock AIME 2024-202584.7%37.8%
LiveBench Math90.2%69.3%
LMArena Math14501412
MATH Level 595.9%78.6%
FrontierMath (Tiers 1-3)24.6%—
FrontierMath Tier 40%—
Omni-MATH41.6%—
FrontierMath (Feb 2025 set)14.1%—
FrontierMath Tier 4 (v1)4.2%—

Knowledge Gemini 2.5 Pro leads

Gemini 2.5 Pro: 56.0 (#46), GPT-4.5: 32.5 (#211)

Knowledge benchmarks
BenchmarkGemini 2.5 ProGPT-4.5
GPQA Diamond85.3%68.7%
Humanity's Last Exam21.6%5.4%
Confabulations10.6%13.6%
LMArena Expert14521394
MMLU-Pro86.3%—
Vectara Hallucination Rate7%—
GPQA (HELM)74.9%—

Multimodal Gemini 2.5 Pro leads

Gemini 2.5 Pro: 45.2 (#18), GPT-4.5: 37.6 (#71)

Multimodal benchmarks
BenchmarkGemini 2.5 ProGPT-4.5
LMArena Vision12631195
VPCT48%45%
GeoBench86%—
LMArena Document1421—
SpatialViz-Bench44.7%—

Multilingual Gemini 2.5 Pro leads

Gemini 2.5 Pro: 55.3 (#31), GPT-4.5: 52.5 (#83)

Multilingual benchmarks
BenchmarkGemini 2.5 ProGPT-4.5
LMArena Non-English14511413
LMArena Chinese15071421
LMArena French14721418
LMArena German14871457
LMArena Japanese14611416
LMArena Korean14341392
LMArena Russian14611419
LMArena Spanish1473—

Instruction Following Gemini 2.5 Pro leads

Gemini 2.5 Pro: 75.0 (#75), GPT-4.5: 72.6 (#134)

Instruction Following benchmarks
BenchmarkGemini 2.5 ProGPT-4.5
LiveBench Instruction Following80.6%72.3%
LMArena Instruction Following14371404
IFEval84%—

Long Context Gemini 2.5 Pro leads

Gemini 2.5 Pro: 59.8 (#5), GPT-4.5: 40.4 (#155)

Long Context benchmarks
BenchmarkGemini 2.5 ProGPT-4.5
Fiction.LiveBench91.7%63.9%
LMArena Longer Query14491406

Writing & Preference Gemini 2.5 Pro leads

Gemini 2.5 Pro: 63.7 (#62), GPT-4.5: 56.9 (#134)

Writing & Preference benchmarks
BenchmarkGemini 2.5 ProGPT-4.5
LMArena Text14581417
LMArena Creative Writing14541394
Short-Story Creative Writing83.8%75.6%
EQ-Bench Creative Writing14211258
LMArena Multi-Turn14531444
LiveBench Language67.8%61.5%
WildBench85.7%—

Frequently asked questions

Is Gemini 2.5 Pro better than GPT-4.5?

Gemini 2.5 Pro is the stronger model overall, scoring 45.0 to 37.2 on the Noometry Index.

Is Gemini 2.5 Pro or GPT-4.5 better for coding?

They score almost the same on coding (42.4 vs 42.2); test both on your own repository before choosing.

How many benchmarks do Gemini 2.5 Pro and GPT-4.5 share?

41 benchmarks have published results for both models. Gemini 2.5 Pro has 78 scored results on Noometry and GPT-4.5 has 42.

Related comparisons

Go deeper