Google, proprietary
Gemini 2.5 Pro
Gemini 2.5 Pro by Google ranks 75th of 354 ranked models on the Noometry Index as of October 2026, with a score of 45.0. Its strongest category is long context, where it ranks 5th. API pricing starts at $1.25 per million input tokens and $10 per million output tokens, with a 1.05M-token context window.
Last verified
Specifications
- Noometry rank
- #75 of 354
- Index score
- 45.0
- Evidence
- Confirmed 78 results
- Provider
Google
- Released
- March 25, 2025
- Weights
- Proprietary
- Reasoning
- Yes
- Context window
- 1.05M
- Max output
- 66K
- Input price
- $1.25 / M
- Output price
- $10 / M
- Blended price
- $3.44 / M
- Output speed
- 5 tokens/s Kagi
- Value
- #171 of 219
- Knowledge cutoff
- January 2025
- Input
- text, image, audio, video, pdf
Category scores
Each category score combines every public result we have in that category.
- Coding 42.4
- Agentic & Tool Use 29.2
- Reasoning 28.8
- Math 32.5
- Knowledge 56.0
- Multimodal 45.2
- Multilingual 55.3
- Instruction Following 75.0
- Long Context 59.8
- Writing & Preference 63.7
| Category | Score | Rank | Results |
|---|---|---|---|
| Coding | 42.4 | #101 | 10 |
| Agentic & Tool Use | 29.2 | #88 | 7 |
| Reasoning | 28.8 | #99 | 12 |
| Math | 32.5 | #213 | 7 |
| Knowledge | 56.0 | #46 | 7 |
| Multimodal | 45.2 | #18 | 3 |
| Multilingual | 55.3 | #31 | 1 |
| Instruction Following | 75.0 | #75 | 3 |
| Long Context | 59.8 | #5 | 2 |
| Writing & Preference | 63.7 | #62 | 7 |
Strengths and weaknesses
Categories where Gemini 2.5 Pro places highest and lowest among the models ranked in each, with its score against that category's median.
Strongest categories
| Category | Score | vs median | Rank |
|---|---|---|---|
| Long Context | 59.8 | +18.8 | #5 of 296, top 2% |
| Multilingual | 55.3 | +7.9 | #31 of 297, top 11% |
| Multimodal | 45.2 | +6.6 | #18 of 128, top 15% |
Weakest categories
| Category | Score | vs median | Rank |
|---|---|---|---|
| Math | 32.5 | −4.1 | #213 of 327, top 66% |
| Agentic & Tool Use | 29.2 | −1.2 | #88 of 154, top 58% |
| Coding | 42.4 | +3.7 | #101 of 340, top 30% |
Closest competitors
The models ranked just above and below Gemini 2.5 Pro. When scores are this close, price and speed are often the better way to choose.
| Model | Rank | Score | Blended $/M | Speed | |
|---|---|---|---|---|---|
| Qwen3.5 Max Preview | #71 | 45.3 | — | — | Compare |
| Qwen3.7 Plus | #72 | 45.3 | $0.70 | — | Compare |
| Hy4 preview | #73 | 45.3 | $1.13 | — | Compare |
| MiMo-V2.5-Pro | #74 | 45.2 | $0.54 | — | Compare |
| GPT-5.4 mini | #76 | 45.0 | $1.69 | 10 | Compare |
| Amazon Nova Experimental Chat 26 02 10 | #77 | 44.5 | — | — | Compare |
| DeepSeek-V3.2-Exp | #78 | 44.3 | $0.29 | 16 | Compare |
| Hy3 | #79 | 44.2 | $0.14 | — | Compare |
Sponsored placements are available on pages like this one. Advertise on Noometry
Benchmark results
Every published result we track, with its source. Bold rows are the ones used for ranking; where several exist we prefer independent runs over self-reported numbers.
Coding
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| SWE-bench Verified | 57.6% | #30 of 32, top 94% | Epoch AI | 2026-02-13 | |
| SWE-bench Verified (bash only) | 53.6% | #27 of 39, top 70% | SWE-bench | 2025-07-26 | |
| Aider Polyglot | 72.9% | Epoch AI | |||
| Aider Polyglot | 76.9% | Epoch AI | |||
| Aider Polyglot | 79.1% | Epoch AI | |||
| Aider Polyglot | 72.9% | Epoch AI | |||
| Aider Polyglot | 83.1% | #3 of 44, top 7% | 32K | Epoch AI | |
| LMArena WebDev | 1227 | #107 of 113, top 95% | LMArena | 2026-10-08 | |
| SciCode | 42.8% | #68 of 121, top 57% | Epoch AI | ||
| GSO | 3.9% | #25 of 31, top 81% | Epoch AI | ||
| GSO | 3.9% | #25 of 31, top 81% | Epoch AI | ||
| WeirdML | 54% | #41 of 119, top 35% | 16K | Epoch AI | |
| LiveBench Coding | 85.9% | Best of 39 | Epoch AI | ||
| LMArena Coding | 1452 | #85 of 294, top 29% | LMArena | 2026-10-08 | |
| CadEval | 64% | #2 of 14, top 15% | Epoch AI | ||
| ALE-Bench | 785.52 | #61 of 105, top 59% | 32K | Epoch AI | |
| AlgoTune | 1.51 | #11 of 18, top 62% | Epoch AI |
Agentic & Tool Use
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| Terminal-Bench | 32.6% | #31 of 41, top 76% | Epoch AI | ||
| GDPval | 23.3% | #9 of 11, top 82% | Epoch AI | ||
| Remote Labor Index | 0.8% | #14 of 14, top 100% | Epoch AI | ||
| TheAgentCompany | 30.3% | #5 of 14, top 36% | Epoch AI | ||
| τ²-bench Banking | 13.7% | #23 of 26, top 89% | high | τ²-bench | 2026-05-05 |
| DeepResearch Bench | 41.5% | Epoch AI | |||
| DeepResearch Bench | 42.8% | #18 of 24, top 75% | Epoch AI | ||
| BALROG | 43.3% | #11 of 35, top 32% | Epoch AI | ||
| LMArena Search | 1142 | #26 of 32, top 82% | LMArena | 2026-08-24 | |
| METR Time Horizons | 55.4% | #21 of 32, top 66% | Epoch AI | ||
| Vending-Bench 2 | 573.64 | #48 of 60, top 80% | Epoch AI |
Reasoning
Math
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| FrontierMath (Tiers 1-3) | 24.6% | #65 of 81, top 81% | Epoch AI | 2026-06-11 | |
| FrontierMath Tier 4 | 0% | #61 of 63, top 97% | Epoch AI | 2026-06-11 | |
| OTIS Mock AIME 2024-2025 | 84.7% | #70 of 173, top 41% | Epoch AI | 2025-11-16 | |
| Omni-MATH | 41.6% | #25 of 57, top 44% | HELM Capabilities | ||
| LiveBench Math | 90.2% | #2 of 39, top 6% | Epoch AI | ||
| LMArena Math | 1450 | #62 of 285, top 22% | LMArena | 2026-10-08 | |
| MATH Level 5 | 95.9% | #10 of 79, top 13% | Epoch AI | 2025-05-08 | |
| MATH Level 5 | 95.6% | Epoch AI | 2025-05-07 | ||
| FrontierMath (Feb 2025 set) | 10.3% | Epoch AI | 2025-07-03 | ||
| FrontierMath (Feb 2025 set) | 14.1% | #35 of 68, top 52% | Epoch AI | 2025-11-24 | |
| FrontierMath Tier 4 (v1) | 4.2% | #32 of 55, top 59% | Epoch AI | 2025-07-03 | |
| FrontierMath Tier 4 (v1) | 2.1% | Epoch AI | 2025-07-03 |
Knowledge
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| GPQA Diamond | 85.3% | #64 of 186, top 35% | Epoch AI | 2025-11-16 | |
| GPQA Diamond | 83.8% | Epoch AI | 2025-03-31 | ||
| GPQA Diamond | 66.7% | Epoch AI | 2025-06-03 | ||
| GPQA Diamond | 84.8% | Epoch AI | 2025-06-05 | ||
| Humanity's Last Exam | 18.2% | Epoch AI | |||
| Humanity's Last Exam | 17.8% | Epoch AI | |||
| Humanity's Last Exam | 21.6% | #17 of 41, top 42% | Epoch AI | ||
| MMLU-Pro | 86.3% | #4 of 58, top 7% | HELM Capabilities | ||
| Confabulations (lower is better) | 10.6% | #2 of 51, top 4% | Lech Mazur benchmarks | ||
| Confabulations (lower is better) | 10.8% | Lech Mazur benchmarks | |||
| Confabulations (lower is better) | 12.4% | Lech Mazur benchmarks | |||
| Vectara Hallucination Rate (lower is better) | 7% | #28 of 96, top 30% | Vectara Hallucination Leaderboard | ||
| GPQA (HELM) | 74.9% | #5 of 57, top 9% | HELM Capabilities | ||
| LMArena Expert | 1452 | #69 of 273, top 26% | LMArena | 2026-10-08 |
Multimodal
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Vision | 1263 | #43 of 122, top 36% | LMArena | 2026-10-09 | |
| GeoBench | 86% | #2 of 25, top 8% | Epoch AI | ||
| GeoBench | 81% | Epoch AI | |||
| VPCT | 48% | #8 of 24, top 34% | Epoch AI | ||
| VPCT | 46.4% | Epoch AI | |||
| VPCT | 40.5% | Epoch AI | |||
| LMArena Document | 1421 | #31 of 38, top 82% | LMArena | 2026-09-13 | |
| SpatialViz-Bench | 44.7% | Best of 8 | Epoch AI |
Multilingual
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Non-English | 1451 | #32 of 297, top 11% | LMArena | 2026-10-08 | |
| LMArena Chinese | 1507 | #40 of 285, top 15% | LMArena | 2026-10-08 | |
| LMArena French | 1472 | #36 of 223, top 17% | LMArena | 2026-10-08 | |
| LMArena German | 1487 | #16 of 231, top 7% | LMArena | 2026-10-08 | |
| LMArena Japanese | 1461 | #20 of 211, top 10% | LMArena | 2026-10-08 | |
| LMArena Korean | 1434 | #27 of 213, top 13% | LMArena | 2026-10-08 | |
| LMArena Russian | 1461 | #32 of 283, top 12% | LMArena | 2026-10-08 | |
| LMArena Spanish | 1473 | #17 of 226, top 8% | LMArena | 2026-10-08 |
Instruction Following
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LiveBench Instruction Following | 80.6% | #10 of 39, top 26% | Epoch AI | ||
| IFEval | 84% | #24 of 57, top 43% | HELM Capabilities | ||
| LMArena Instruction Following | 1437 | #54 of 298, top 19% | LMArena | 2026-10-08 |
Long Context
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| Fiction.LiveBench | 66.7% | Epoch AI | |||
| Fiction.LiveBench | 91.7% | #5 of 47, top 11% | Epoch AI | ||
| Fiction.LiveBench | 66.7% | Epoch AI | |||
| Fiction.LiveBench | 66.7% | Epoch AI | |||
| LMArena Longer Query | 1449 | #54 of 291, top 19% | LMArena | 2026-10-08 |
Writing & Preference
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Text | 1458 | #36 of 297, top 13% | LMArena | 2026-10-08 | |
| LMArena Creative Writing | 1454 | #26 of 295, top 9% | LMArena | 2026-10-08 | |
| Short-Story Creative Writing | 80.9% | Epoch AI | |||
| Short-Story Creative Writing | 80.5% | Epoch AI | |||
| Short-Story Creative Writing | 83.8% | #6 of 39, top 16% | Epoch AI | ||
| EQ-Bench Creative Writing | 1421 | #65 of 115, top 57% | EQ-Bench | ||
| EQ-Bench Creative Writing | 1396 | EQ-Bench | |||
| WildBench | 85.7% | #7 of 57, top 13% | HELM Capabilities | ||
| LMArena Multi-Turn | 1453 | #50 of 295, top 17% | LMArena | 2026-10-08 | |
| LiveBench Language | 67.8% | #2 of 39, top 6% | Epoch AI |
API pricing by provider
| Route | Input $/M | Output $/M | Cached input $/M | Checked |
|---|---|---|---|---|
| $1.25 | $10 | $0.13 | 2026-10-10 | |
| openrouter | $1.25 | $10 | $0.13 | 2026-10-10 |
| vertex | $1.25 | $10 | $0.13 | 2026-10-10 |
Compare Gemini 2.5 Pro
- Gemini 2.5 Pro vs Gemini 2.0 Pro
- Gemini 2.5 Pro vs MiMo-V2.5-Pro
- Gemini 2.5 Pro vs GPT-5.4 mini
- Gemini 2.5 Pro vs Hy4 preview
- Gemini 2.5 Pro vs Amazon Nova Experimental Chat 26 02 10
- Gemini 2.5 Pro vs Qwen3.7 Plus
- Gemini 2.5 Pro vs DeepSeek-V3.2-Exp
- Gemini 2.5 Pro vs GPT-6 Astra
- Gemini 2.5 Pro vs Claude Fable 5.1
- Gemini 2.5 Pro vs Kimi K3
- Gemini 2.5 Pro vs Grok 4.6
- Gemini 2.5 Pro vs Qwen3.8 Max
- Gemini 2.5 Pro vs GLM-5.3
- Gemini 2.5 Pro vs Muse Spark 1.3
Other Google models
Frequently asked questions
How good is Gemini 2.5 Pro?
Gemini 2.5 Pro by Google ranks 75th of 354 ranked models on the Noometry Index as of October 2026, with a score of 45.0. Its strongest category is long context, where it ranks 5th. API pricing starts at $1.25 per million input tokens and $10 per million output tokens, with a 1.05M-token context window.
How much does Gemini 2.5 Pro cost?
Gemini 2.5 Pro costs $1.25 per million input tokens and $10 per million output tokens on Google's own API, with cached input at $0.13.
What is Gemini 2.5 Pro's context window?
Gemini 2.5 Pro accepts up to 1.05M tokens of input and can write up to 66K tokens in one response.
Is Gemini 2.5 Pro open source?
No. Gemini 2.5 Pro is proprietary and available only through Google's API and partner platforms.
How fast is Gemini 2.5 Pro?
Gemini 2.5 Pro generated about 5 output tokens per second in the Kagi LLM Benchmark's timed runs. Speed varies by provider, load and reasoning effort.
What are Gemini 2.5 Pro's strengths and weaknesses?
Relative to other ranked models, Gemini 2.5 Pro places best in long context, multilingual, multimodal and lowest in math, agentic & tool use, coding.
What is Gemini 2.5 Pro best at?
Its best category is long context, where it ranks 5th on Noometry.