Model comparison

Llama 3.2 90B vs Llama 4 Scout

Llama 3.2 90B and Llama 4 Scout score almost the same on the Noometry Index (27.5 vs 27.7), so choose on price, context window or the category you care about most.

Last verified . 5 shared benchmarks.

Llama 3.2 90B Meta

27.5

Rank #331 Confirmed

Llama 4 Scout Meta

27.7

Rank #330 Confirmed

Summary

  • They share 5 benchmarks with published results for both. Llama 3.2 90B scores higher in 2 categories and Llama 4 Scout in 3 categories; 5 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Llama 3.2 90B leads 21.7 to 9.1.
  • The biggest single-benchmark swing is MATH Level 5: 39.4% for Llama 3.2 90B and 62.3% for Llama 4 Scout.

Side by side

Llama 3.2 90B and Llama 4 Scout specifications
Llama 3.2 90BLlama 4 Scout
ProviderMetaMeta
Noometry Index27.527.7
Released2024-09-242025-04-05
WeightsOpenOpen
Context window—128K
Max output—4K
Input $ / M tokens—$0.10
Output $ / M tokens—$0.30
Results tracked943

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Not comparable

Llama 3.2 90B: —, Llama 4 Scout: 20.2 (#339)

Coding benchmarks
BenchmarkLlama 3.2 90BLlama 4 Scout
SWE-bench Verified (bash only)—9.1%
SciCode—17%
LMArena Coding—1286
BigCodeBench Complete—43.1%

Agentic & Tool Use Llama 3.2 90B leads

Llama 3.2 90B: 30.0 (#80), Llama 4 Scout: 24.6 (#119)

Agentic & Tool Use benchmarks
BenchmarkLlama 3.2 90BLlama 4 Scout
Berkeley Function Calling Leaderboard—28.1%
BALROG27.3%—

Reasoning Llama 3.2 90B leads

Llama 3.2 90B: 21.7 (#217), Llama 4 Scout: 9.1 (#345)

Reasoning benchmarks
BenchmarkLlama 3.2 90BLlama 4 Scout
Epoch Capabilities Index125.5129.64
ARC-AGI-2—0%
Kagi LLM Benchmark—36.9%
ARC-AGI-1—0.5%
CritPt—0%
EnigmaEval0.4%—
LMArena Hard Prompts—1266
DTBench—57.9%
LMCA—12%
ForecastBench—57.5

Math Llama 4 Scout leads

Llama 3.2 90B: 11.1 (#308), Llama 4 Scout: 19.6 (#286)

Math benchmarks
BenchmarkLlama 3.2 90BLlama 4 Scout
OTIS Mock AIME 2024-20252.6%7.8%
MATH Level 539.4%62.3%
Omni-MATH—37.3%
LMArena Math—1287
FrontierMath (Feb 2025 set)—0%

Knowledge Llama 4 Scout leads

Llama 3.2 90B: 21.7 (#274), Llama 4 Scout: 31.9 (#217)

Knowledge benchmarks
BenchmarkLlama 3.2 90BLlama 4 Scout
GPQA Diamond41%51.8%
MMLU-Pro—74.2%
Vectara Hallucination Rate—7.7%
GPQA (HELM)—50.7%
LMArena Expert—1235
MMLU80.3%—

Multimodal Llama 4 Scout leads

Llama 3.2 90B: 25.4 (#124), Llama 4 Scout: 32.2 (#102)

Multimodal benchmarks
BenchmarkLlama 3.2 90BLlama 4 Scout
LMArena Vision10001118
GeoBench52%—
SpatialViz-Bench—34.2%

Multilingual Not comparable

Llama 3.2 90B: —, Llama 4 Scout: 41.0 (#212)

Multilingual benchmarks
BenchmarkLlama 3.2 90BLlama 4 Scout
LMArena Non-English—1252
LMArena Chinese—1255
LMArena French—1282
LMArena German—1272
LMArena Japanese—1206
LMArena Korean—1207
LMArena Russian—1263
LMArena Spanish—1278

Instruction Following Not comparable

Llama 3.2 90B: —, Llama 4 Scout: 65.8 (#217)

Instruction Following benchmarks
BenchmarkLlama 3.2 90BLlama 4 Scout
IFEval—81.8%
LMArena Instruction Following—1248

Long Context Not comparable

Llama 3.2 90B: —, Llama 4 Scout: 27.5 (#294)

Long Context benchmarks
BenchmarkLlama 3.2 90BLlama 4 Scout
Fiction.LiveBench—36%
LMArena Longer Query—1265

Writing & Preference Not comparable

Llama 3.2 90B: —, Llama 4 Scout: 37.0 (#261)

Writing & Preference benchmarks
BenchmarkLlama 3.2 90BLlama 4 Scout
LMArena Text—1279
LMArena Creative Writing—1249
EQ-Bench Creative Writing—783
WildBench—78%
LMArena Multi-Turn—1280

Frequently asked questions

Is Llama 3.2 90B better than Llama 4 Scout?

Llama 3.2 90B and Llama 4 Scout score almost the same on the Noometry Index (27.5 vs 27.7), so choose on price, context window or the category you care about most.

How many benchmarks do Llama 3.2 90B and Llama 4 Scout share?

5 benchmarks have published results for both models. Llama 3.2 90B has 9 scored results on Noometry and Llama 4 Scout has 43.

Related comparisons

Go deeper