Google, proprietary
Gemini 2.0 Flash (Feb 2025)
Gemini 2.0 Flash (Feb 2025) by Google ranks 228th of 354 ranked models on the Noometry Index as of October 2026, with a score of 35.1. Its strongest category is multimodal, where it ranks 79th.
Last verified
Specifications
- Noometry rank
- #228 of 354
- Index score
- 35.1
- Evidence
- Confirmed 54 results
- Provider
Google
- Released
- December 6, 2024
- Weights
- Proprietary
- Reasoning
- Unknown
- Context window
- —
- Max output
- —
- Input price
- Not listed
- Output price
- Not listed
- Blended price
- Not listed
- Output speed
- 92 tokens/s Kagi
- Value
- Not ranked
- Knowledge cutoff
- Unknown
Category scores
Each category score combines every public result we have in that category.
- Coding 28.4
- Agentic & Tool Use 28.1
- Reasoning 15.2
- Math 37.9
- Knowledge 32.0
- Multimodal 36.5
- Multilingual 47.4
- Instruction Following 74.4
- Long Context 38.1
- Writing & Preference 49.5
| Category | Score | Rank | Results |
|---|---|---|---|
| Coding | 28.4 | #315 | 8 |
| Agentic & Tool Use | 28.1 | #92 | 1 |
| Reasoning | 15.2 | #318 | 8 |
| Math | 37.9 | #146 | 5 |
| Knowledge | 32.0 | #213 | 6 |
| Multimodal | 36.5 | #79 | 2 |
| Multilingual | 47.4 | #149 | 1 |
| Instruction Following | 74.4 | #97 | 3 |
| Long Context | 38.1 | #203 | 2 |
| Writing & Preference | 49.5 | #190 | 7 |
Strengths and weaknesses
Categories where Gemini 2.0 Flash (Feb 2025) places highest and lowest among the models ranked in each, with its score against that category's median.
Strongest categories
| Category | Score | vs median | Rank |
|---|---|---|---|
| Instruction Following | 74.4 | +3.1 | #97 of 305, top 32% |
| Math | 37.9 | +1.3 | #146 of 327, top 45% |
| Multilingual | 47.4 | −0.0 | #149 of 297, top 51% |
Weakest categories
| Category | Score | vs median | Rank |
|---|---|---|---|
| Coding | 28.4 | −10.3 | #315 of 340, top 93% |
| Reasoning | 15.2 | −8.4 | #318 of 350, top 91% |
| Long Context | 38.1 | −2.8 | #203 of 296, top 69% |
Closest competitors
The models ranked just above and below Gemini 2.0 Flash (Feb 2025). When scores are this close, price and speed are often the better way to choose.
| Model | Rank | Score | Blended $/M | Speed | |
|---|---|---|---|---|---|
| Llama 3.1 Tulu 3 8b | #224 | 35.7 | — | — | Compare |
| Qwen3 14B | #225 | 35.5 | $0.61 | 79 | Compare |
| DeepSeek-R1-Distill-Qwen-32B | #226 | 35.5 | — | — | Compare |
| Magistral Medium | #227 | 35.2 | $2.75 | 0 | Compare |
| C4ai Aya Expanse 8b | #229 | 34.9 | — | — | Compare |
| Qwen Max | #230 | 34.7 | $2.80 | — | Compare |
| Claude 3.5 Sonnet | #231 | 34.6 | — | — | Compare |
| Qwen3 Coder Next | #232 | 34.3 | $0.29 | — | Compare |
Sponsored placements are available on pages like this one. Advertise on Noometry
Benchmark results
Every published result we track, with its source. Bold rows are the ones used for ranking; where several exist we prefer independent runs over self-reported numbers.
Coding
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| SWE-bench Verified (bash only) | 13.5% | #37 of 39, top 95% | SWE-bench | 2025-07-26 | |
| Aider Polyglot | 18.2% | Epoch AI | |||
| Aider Polyglot | 22.2% | Epoch AI | |||
| Aider Polyglot | 38.2% | #29 of 44, top 66% | Epoch AI | ||
| WeirdML | 25.8% | #98 of 119, top 83% | Epoch AI | ||
| BigCodeBench Instruct | 45.9% | #16 of 64, top 25% | BigCodeBench | 2025-02-05 | |
| LiveBench Coding | 53.5% | Epoch AI | |||
| LiveBench Coding | 53.9% | Epoch AI | |||
| LiveBench Coding | 54.4% | Epoch AI | |||
| LiveBench Coding | 63.4% | #13 of 39, top 34% | Epoch AI | ||
| LMArena Coding | 1350 | #173 of 294, top 59% | LMArena | 2026-10-08 | |
| BigCodeBench Complete | 59.9% | #4 of 66, top 7% | BigCodeBench | 2025-02-05 | |
| CadEval | 30% | #11 of 14, top 79% | Epoch AI |
Agentic & Tool Use
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| TheAgentCompany | 11.4% | #7 of 14, top 50% | Epoch AI |
Reasoning
Math
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| OTIS Mock AIME 2024-2025 | 57.8% | #106 of 173, top 62% | Epoch AI | 2025-03-07 | |
| OTIS Mock AIME 2024-2025 | 31.1% | Epoch AI | 2025-02-25 | ||
| Omni-MATH | 45.9% | #22 of 57, top 39% | HELM Capabilities | ||
| LiveBench Math | 65.6% | Epoch AI | |||
| LiveBench Math | 75.8% | #8 of 39, top 21% | Epoch AI | ||
| LiveBench Math | 72.4% | Epoch AI | |||
| LiveBench Math | 60.4% | Epoch AI | |||
| LMArena Math | 1352 | #165 of 285, top 58% | LMArena | 2026-10-08 | |
| MATH Level 5 | 82.2% | #24 of 79, top 31% | Epoch AI | 2025-02-06 | |
| FrontierMath (Feb 2025 set) | 1.7% | #56 of 68, top 83% | Epoch AI | 2025-03-09 |
Knowledge
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| GPQA Diamond | 64.1% | #110 of 186, top 60% | Epoch AI | 2025-02-06 | |
| GPQA Diamond | 57.1% | Epoch AI | 2025-02-06 | ||
| Humanity's Last Exam | 6.6% | #32 of 41, top 79% | Epoch AI | ||
| MMLU-Pro | 73.7% | #30 of 58, top 52% | HELM Capabilities | ||
| Confabulations (lower is better) | 26.9% | Lech Mazur benchmarks | |||
| Confabulations (lower is better) | 12.4% | #8 of 51, top 16% | Lech Mazur benchmarks | ||
| GPQA (HELM) | 55.6% | #27 of 57, top 48% | HELM Capabilities | ||
| LMArena Expert | 1339 | #158 of 273, top 58% | LMArena | 2026-10-08 | |
| MMLU | 79.7% | #19 of 81, top 24% | Epoch AI |
Multimodal
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Vision | 1158 | #97 of 122, top 80% | LMArena | 2026-10-09 | |
| GeoBench | 77% | #6 of 25, top 24% | Epoch AI |
Multilingual
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Non-English | 1342 | #149 of 297, top 51% | LMArena | 2026-10-08 | |
| LMArena Chinese | 1373 | #154 of 285, top 55% | LMArena | 2026-10-08 | |
| LMArena French | 1391 | #125 of 223, top 57% | LMArena | 2026-10-08 | |
| LMArena German | 1353 | #127 of 231, top 55% | LMArena | 2026-10-08 | |
| LMArena Japanese | 1294 | #124 of 211, top 59% | LMArena | 2026-10-08 | |
| LMArena Korean | 1313 | #123 of 213, top 58% | LMArena | 2026-10-08 | |
| LMArena Russian | 1351 | #147 of 283, top 52% | LMArena | 2026-10-08 | |
| LMArena Spanish | 1363 | #134 of 226, top 60% | LMArena | 2026-10-08 |
Instruction Following
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LiveBench Instruction Following | 85.8% | #2 of 39, top 6% | Epoch AI | ||
| LiveBench Instruction Following | 81.9% | Epoch AI | |||
| LiveBench Instruction Following | 77.3% | Epoch AI | |||
| LiveBench Instruction Following | 82.5% | Epoch AI | |||
| IFEval | 84.1% | #22 of 57, top 39% | HELM Capabilities | ||
| LMArena Instruction Following | 1336 | #152 of 298, top 52% | LMArena | 2026-10-08 |
Long Context
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| Fiction.LiveBench | 61.1% | #29 of 47, top 62% | Epoch AI | ||
| Fiction.LiveBench | 52.8% | Epoch AI | |||
| LMArena Longer Query | 1344 | #155 of 291, top 54% | LMArena | 2026-10-08 |
Writing & Preference
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Text | 1354 | #155 of 297, top 53% | LMArena | 2026-10-08 | |
| LMArena Creative Writing | 1340 | #138 of 295, top 47% | LMArena | 2026-10-08 | |
| Short-Story Creative Writing | 71.5% | Epoch AI | |||
| Short-Story Creative Writing | 73.8% | #26 of 39, top 67% | Epoch AI | ||
| EQ-Bench Creative Writing | 1128 | #91 of 115, top 80% | EQ-Bench | ||
| WildBench | 80% | #30 of 57, top 53% | HELM Capabilities | ||
| LMArena Multi-Turn | 1350 | #155 of 295, top 53% | LMArena | 2026-10-08 | |
| LiveBench Language | 40.7% | Epoch AI | |||
| LiveBench Language | 51.3% | #9 of 39, top 24% | Epoch AI | ||
| LiveBench Language | 38.2% | Epoch AI | |||
| LiveBench Language | 42.2% | Epoch AI |
Compare Gemini 2.0 Flash (Feb 2025)
- Gemini 2.0 Flash (Feb 2025) vs Magistral Medium
- Gemini 2.0 Flash (Feb 2025) vs C4ai Aya Expanse 8b
- Gemini 2.0 Flash (Feb 2025) vs DeepSeek-R1-Distill-Qwen-32B
- Gemini 2.0 Flash (Feb 2025) vs Qwen Max
- Gemini 2.0 Flash (Feb 2025) vs Qwen3 14B
- Gemini 2.0 Flash (Feb 2025) vs Claude 3.5 Sonnet
- Gemini 2.0 Flash (Feb 2025) vs GPT-6 Astra
- Gemini 2.0 Flash (Feb 2025) vs Claude Fable 5.1
- Gemini 2.0 Flash (Feb 2025) vs Kimi K3
- Gemini 2.0 Flash (Feb 2025) vs Grok 4.6
- Gemini 2.0 Flash (Feb 2025) vs Qwen3.8 Max
- Gemini 2.0 Flash (Feb 2025) vs GLM-5.3
- Gemini 2.0 Flash (Feb 2025) vs Muse Spark 1.3
- Gemini 2.0 Flash (Feb 2025) vs DeepSeek V4 Pro
Other Google models
Frequently asked questions
How good is Gemini 2.0 Flash (Feb 2025)?
Gemini 2.0 Flash (Feb 2025) by Google ranks 228th of 354 ranked models on the Noometry Index as of October 2026, with a score of 35.1. Its strongest category is multimodal, where it ranks 79th.
Is Gemini 2.0 Flash (Feb 2025) open source?
No. Gemini 2.0 Flash (Feb 2025) is proprietary and available only through Google's API and partner platforms.
How fast is Gemini 2.0 Flash (Feb 2025)?
Gemini 2.0 Flash (Feb 2025) generated about 92 output tokens per second in the Kagi LLM Benchmark's timed runs. Speed varies by provider, load and reasoning effort.
What are Gemini 2.0 Flash (Feb 2025)'s strengths and weaknesses?
Relative to other ranked models, Gemini 2.0 Flash (Feb 2025) places best in instruction following, math, multilingual and lowest in coding, reasoning, long context.
What is Gemini 2.0 Flash (Feb 2025) best at?
Its best category is multimodal, where it ranks 79th on Noometry.