Meta, open weights

Llama 4 Scout

Llama 4 Scout by Meta ranks 330th of 354 ranked models on the Noometry Index as of October 2026, with a score of 27.7. Its strongest category is multimodal, where it ranks 102nd. API pricing starts at $0.10 per million input tokens and $0.30 per million output tokens, with a 128K-token context window.

Last verified

Specifications

Noometry rank
#330 of 354
Index score
27.7
Evidence
Confirmed 43 results
Provider
Meta
Released
April 5, 2025
Weights
Open weights
Reasoning
No
Context window
128K
Max output
4K
Input price
$0.10 / M
Output price
$0.30 / M
Blended price
$0.15 / M
Output speed
272 tokens/s Kagi
Value
#45 of 219
Knowledge cutoff
August 2024
Input
text, image

Category scores

Each category score combines every public result we have in that category.

Llama 4 Scout category scores
  1. Coding 20.2
  2. Agentic & Tool Use 24.6
  3. Reasoning 9.1
  4. Math 19.6
  5. Knowledge 31.9
  6. Multimodal 32.2
  7. Multilingual 41.0
  8. Instruction Following 65.8
  9. Long Context 27.5
  10. Writing & Preference 37.0
Llama 4 Scout category ranks
CategoryScoreRankResults
Coding20.2#3394
Agentic & Tool Use24.6#1191
Reasoning9.1#3457
Math19.6#2864
Knowledge31.9#2175
Multimodal32.2#1021
Multilingual41.0#2121
Instruction Following65.8#2172
Long Context27.5#2942
Writing & Preference37.0#2615

Strengths and weaknesses

Categories where Llama 4 Scout places highest and lowest among the models ranked in each, with its score against that category's median.

Strongest categories

Llama 4 Scout: strongest categories
CategoryScorevs medianRank
Knowledge31.9−5.4#217 of 314, top 70%
Instruction Following65.8−5.5#217 of 305, top 72%
Multilingual41.0−6.4#212 of 297, top 72%

Weakest categories

Llama 4 Scout: weakest categories
CategoryScorevs medianRank
Coding20.2−18.5#339 of 340, top 100%
Long Context27.5−13.4#294 of 296, top 100%
Reasoning9.1−14.5#345 of 350, top 99%

Closest competitors

The models ranked just above and below Llama 4 Scout. When scores are this close, price and speed are often the better way to choose.

Models ranked closest to Llama 4 Scout
ModelRankScoreBlended $/MSpeed
Gemma 3 4B#32628.1$0.0572Compare
GPT-4.1 nano#32727.9$0.18135Compare
Phi 3 Mini 4k Instruct#32827.9——Compare
Yi-34B#32927.8——Compare
Llama 3.2 90B#33127.5——Compare
Gemini 1.0 Pro#33227.3——Compare
Mixtral 8x22B#33327.1$3—Compare
Mixtral 8x7B#33427.1$0.70—Compare

Sponsored placements are available on pages like this one. Advertise on Noometry

Benchmark results

Every published result we track, with its source. Bold rows are the ones used for ranking; where several exist we prefer independent runs over self-reported numbers.

Coding

Llama 4 Scout Coding benchmark results
BenchmarkScorePositionSettingSourceDate
SWE-bench Verified (bash only)9.1%#38 of 39, top 98%SWE-bench2025-07-20
SciCode17%#119 of 121, top 99%Epoch AI
LMArena Coding1286#209 of 294, top 72%LMArena2026-10-08
BigCodeBench Complete43.1%#47 of 66, top 72%BigCodeBench2025-04-05

Agentic & Tool Use

Llama 4 Scout Agentic & Tool Use benchmark results
BenchmarkScorePositionSettingSourceDate
Berkeley Function Calling Leaderboard28.1%#37 of 49, top 76%fcBerkeley Function Calling Leaderboard

Reasoning

Llama 4 Scout Reasoning benchmark results
BenchmarkScorePositionSettingSourceDate
ARC-AGI-20%#81 of 83, top 98%Epoch AI
Kagi LLM Benchmark36.9%#86 of 99, top 87%Kagi LLM Benchmark
ARC-AGI-10.5%#82 of 83, top 99%Epoch AI
CritPt0%#120 of 134, top 90%Epoch AI
LMArena Hard Prompts1266#211 of 297, top 72%LMArena2026-10-08
DTBench57.9%#120 of 151, top 80%Epoch AI
LMCA12%#109 of 125, top 88%Epoch AI
Epoch Capabilities Index129.64#139 of 213, top 66%Epoch AI2025-04-05
ForecastBench57.5#57 of 72, top 80%Epoch AI

Math

Llama 4 Scout Math benchmark results
BenchmarkScorePositionSettingSourceDate
OTIS Mock AIME 2024-20257.8%#138 of 173, top 80%Epoch AI2025-04-08
Omni-MATH37.3%#30 of 57, top 53%HELM Capabilities
LMArena Math1287#188 of 285, top 66%LMArena2026-10-08
MATH Level 562.3%#38 of 79, top 49%Epoch AI2025-04-08
FrontierMath (Feb 2025 set)0%#68 of 68, top 100%Epoch AI2025-04-08

Knowledge

Llama 4 Scout Knowledge benchmark results
BenchmarkScorePositionSettingSourceDate
GPQA Diamond51.8%#125 of 186, top 68%Epoch AI2025-04-08
MMLU-Pro74.2%#27 of 58, top 47%HELM Capabilities
Vectara Hallucination Rate (lower is better)7.7%#34 of 96, top 36%Vectara Hallucination Leaderboard
GPQA (HELM)50.7%#34 of 57, top 60%HELM Capabilities
LMArena Expert1235#205 of 273, top 76%LMArena2026-10-08

Multimodal

Llama 4 Scout Multimodal benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Vision1118#105 of 122, top 87%LMArena2026-10-09
SpatialViz-Bench34.2%#4 of 8, top 50%Epoch AI

Multilingual

Llama 4 Scout Multilingual benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Non-English1252#212 of 297, top 72%LMArena2026-10-08
LMArena Chinese1255#204 of 285, top 72%LMArena2026-10-08
LMArena French1282#169 of 223, top 76%LMArena2026-10-08
LMArena German1272#165 of 231, top 72%LMArena2026-10-08
LMArena Japanese1206#157 of 211, top 75%LMArena2026-10-08
LMArena Korean1207#156 of 213, top 74%LMArena2026-10-08
LMArena Russian1263#202 of 283, top 72%LMArena2026-10-08
LMArena Spanish1278#170 of 226, top 76%LMArena2026-10-08

Instruction Following

Llama 4 Scout Instruction Following benchmark results
BenchmarkScorePositionSettingSourceDate
IFEval81.8%#34 of 57, top 60%HELM Capabilities
LMArena Instruction Following1248#215 of 298, top 73%LMArena2026-10-08

Long Context

Llama 4 Scout Long Context benchmark results
BenchmarkScorePositionSettingSourceDate
Fiction.LiveBench36%#44 of 47, top 94%Epoch AI
LMArena Longer Query1265#213 of 291, top 74%LMArena2026-10-08

Writing & Preference

Llama 4 Scout Writing & Preference benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Text1279#210 of 297, top 71%LMArena2026-10-08
LMArena Creative Writing1249#206 of 295, top 70%LMArena2026-10-08
EQ-Bench Creative Writing783#105 of 115, top 92%EQ-Bench
WildBench78%#41 of 57, top 72%HELM Capabilities
LMArena Multi-Turn1280#200 of 295, top 68%LMArena2026-10-08

API pricing by provider

Llama 4 Scout API prices
RouteInput $/MOutput $/MCached input $/MChecked
azure$0.20$0.78—2026-10-10
bedrock$0.17$0.66—2026-10-10
deepinfra$0.10$0.30—2026-10-10
openrouter$0.10$0.30—2026-10-10

Compare Llama 4 Scout

Other Meta models

Frequently asked questions

How good is Llama 4 Scout?

Llama 4 Scout by Meta ranks 330th of 354 ranked models on the Noometry Index as of October 2026, with a score of 27.7. Its strongest category is multimodal, where it ranks 102nd. API pricing starts at $0.10 per million input tokens and $0.30 per million output tokens, with a 128K-token context window.

How much does Llama 4 Scout cost?

Llama 4 Scout costs $0.10 per million input tokens and $0.30 per million output tokens on deepinfra.

What is Llama 4 Scout's context window?

Llama 4 Scout accepts up to 128K tokens of input and can write up to 4K tokens in one response.

Is Llama 4 Scout open source?

Yes. Llama 4 Scout's weights are downloadable from Hugging Face (meta-llama/Llama-4-Scout-17B-16E-Instruct); check the license for commercial terms.

How fast is Llama 4 Scout?

Llama 4 Scout generated about 272 output tokens per second in the Kagi LLM Benchmark's timed runs. Speed varies by provider, load and reasoning effort.

What are Llama 4 Scout's strengths and weaknesses?

Relative to other ranked models, Llama 4 Scout places best in knowledge, instruction following, multilingual and lowest in coding, long context, reasoning.

What is Llama 4 Scout best at?

Its best category is multimodal, where it ranks 102nd on Noometry.