Model comparison
Gemini 1.0 Pro vs Llama 4 Scout
Gemini 1.0 Pro and Llama 4 Scout score almost the same on the Noometry Index (27.3 vs 27.7), so choose on price, context window or the category you care about most.
Last verified . 21 shared benchmarks.
Summary
- They share 21 benchmarks with published results for both. Gemini 1.0 Pro scores higher in 3 categories and Llama 4 Scout in 5 categories; 7 gaps are clear of the uncertainty.
- The widest gap is in knowledge, where Llama 4 Scout leads 31.9 to 15.6.
- The biggest single-benchmark swing is MATH Level 5: 11.2% for Gemini 1.0 Pro and 62.3% for Llama 4 Scout.
- Llama 4 Scout has downloadable open weights; the other is API-only.
Side by side
| Gemini 1.0 Pro | Llama 4 Scout | |
|---|---|---|
| Provider | Meta | |
| Noometry Index | 27.3 | 27.7 |
| Released | 2023-12-13 | 2025-04-05 |
| Weights | Proprietary | Open |
| Context window | — | 128K |
| Max output | — | 4K |
| Input $ / M tokens | — | $0.10 |
| Output $ / M tokens | — | $0.30 |
| Results tracked | 24 | 43 |
Sponsored placements are available on pages like this one. Advertise on Noometry
Category by category
Coding Gemini 1.0 Pro leads
Gemini 1.0 Pro: 32.2 (#275), Llama 4 Scout: 20.2 (#339)
| Benchmark | Gemini 1.0 Pro | Llama 4 Scout |
|---|---|---|
| LMArena Coding | 1108 | 1286 |
| SWE-bench Verified (bash only) | — | 9.1% |
| SciCode | — | 17% |
| BigCodeBench Complete | — | 43.1% |
| HumanEval+ | 55.5% | — |
| MBPP+ | 61.4% | — |
Agentic & Tool Use Not comparable
Gemini 1.0 Pro: —, Llama 4 Scout: 24.6 (#119)
| Benchmark | Gemini 1.0 Pro | Llama 4 Scout |
|---|---|---|
| Berkeley Function Calling Leaderboard | — | 28.1% |
Reasoning Gemini 1.0 Pro leads
Gemini 1.0 Pro: 17.1 (#296), Llama 4 Scout: 9.1 (#345)
| Benchmark | Gemini 1.0 Pro | Llama 4 Scout |
|---|---|---|
| LMArena Hard Prompts | 1109 | 1266 |
| DTBench | 45.9% | 57.9% |
| Epoch Capabilities Index | 117.04 | 129.64 |
| ARC-AGI-2 | — | 0% |
| Kagi LLM Benchmark | — | 36.9% |
| ARC-AGI-1 | — | 0.5% |
| CritPt | — | 0% |
| LMCA | — | 12% |
| ForecastBench | — | 57.5 |
Math Llama 4 Scout leads
Gemini 1.0 Pro: 9.3 (#321), Llama 4 Scout: 19.6 (#286)
| Benchmark | Gemini 1.0 Pro | Llama 4 Scout |
|---|---|---|
| OTIS Mock AIME 2024-2025 | 1.1% | 7.8% |
| LMArena Math | 1132 | 1287 |
| MATH Level 5 | 11.2% | 62.3% |
| Omni-MATH | — | 37.3% |
| FrontierMath (Feb 2025 set) | — | 0% |
Knowledge Llama 4 Scout leads
Gemini 1.0 Pro: 15.6 (#291), Llama 4 Scout: 31.9 (#217)
| Benchmark | Gemini 1.0 Pro | Llama 4 Scout |
|---|---|---|
| GPQA Diamond | 34% | 51.8% |
| LMArena Expert | 1059 | 1235 |
| MMLU-Pro | — | 74.2% |
| Vectara Hallucination Rate | — | 7.7% |
| GPQA (HELM) | — | 50.7% |
| MMLU | 70% | — |
Multimodal Not comparable
Gemini 1.0 Pro: —, Llama 4 Scout: 32.2 (#102)
| Benchmark | Gemini 1.0 Pro | Llama 4 Scout |
|---|---|---|
| LMArena Vision | — | 1118 |
| SpatialViz-Bench | — | 34.2% |
Multilingual Llama 4 Scout leads
Gemini 1.0 Pro: 33.4 (#252), Llama 4 Scout: 41.0 (#212)
| Benchmark | Gemini 1.0 Pro | Llama 4 Scout |
|---|---|---|
| LMArena Non-English | 1138 | 1252 |
| LMArena Chinese | 1124 | 1255 |
| LMArena French | 1145 | 1282 |
| LMArena German | 1125 | 1272 |
| LMArena Japanese | 1023 | 1206 |
| LMArena Russian | 1186 | 1263 |
| LMArena Spanish | 1119 | 1278 |
| LMArena Korean | — | 1207 |
Instruction Following Llama 4 Scout leads
Gemini 1.0 Pro: 57.6 (#267), Llama 4 Scout: 65.8 (#217)
| Benchmark | Gemini 1.0 Pro | Llama 4 Scout |
|---|---|---|
| LMArena Instruction Following | 1114 | 1248 |
| IFEval | — | 81.8% |
Long Context Gemini 1.0 Pro leads
Gemini 1.0 Pro: 34.3 (#249), Llama 4 Scout: 27.5 (#294)
| Benchmark | Gemini 1.0 Pro | Llama 4 Scout |
|---|---|---|
| LMArena Longer Query | 1132 | 1265 |
| Fiction.LiveBench | — | 36% |
Writing & Preference Too close to call
Gemini 1.0 Pro: 36.0 (#264), Llama 4 Scout: 37.0 (#261)
| Benchmark | Gemini 1.0 Pro | Llama 4 Scout |
|---|---|---|
| LMArena Text | 1149 | 1279 |
| LMArena Creative Writing | 1131 | 1249 |
| LMArena Multi-Turn | 1139 | 1280 |
| EQ-Bench Creative Writing | — | 783 |
| WildBench | — | 78% |
Frequently asked questions
Is Gemini 1.0 Pro better than Llama 4 Scout?
Gemini 1.0 Pro and Llama 4 Scout score almost the same on the Noometry Index (27.3 vs 27.7), so choose on price, context window or the category you care about most.
Is Gemini 1.0 Pro or Llama 4 Scout better for coding?
Gemini 1.0 Pro scores higher on coding benchmarks: 32.2 versus 20.2 in the Noometry coding category.
How many benchmarks do Gemini 1.0 Pro and Llama 4 Scout share?
21 benchmarks have published results for both models. Gemini 1.0 Pro has 24 scored results on Noometry and Llama 4 Scout has 43.