Model comparison

Llama 4 Maverick vs Mistral Small 3.2

Llama 4 Maverick and Mistral Small 3.2 score almost the same on the Noometry Index (30.9 vs 31.2), so choose on price, context window or the category you care about most.

Last verified . 5 shared benchmarks.

Llama 4 Maverick Meta

30.9

Rank #282 Confirmed

Mistral Small 3.2 Mistral AI

31.2

Rank #280 Confirmed

Summary

  • They share 5 benchmarks with published results for both. Llama 4 Maverick scores higher in 1 category and Mistral Small 3.2 in 3 categories; 3 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Mistral Small 3.2 leads 18.1 to 10.1.
  • The biggest single-benchmark swing is GPQA Diamond: 67% for Llama 4 Maverick and 49.1% for Mistral Small 3.2.
  • Mistral Small 3.2 is cheaper at $0.0938 / $0.25 per million input/output tokens, against $0.19 / $0.65 for Llama 4 Maverick.
  • Mistral Small 3.2 accepts more context: 256K tokens versus 128K.

Side by side

Llama 4 Maverick and Mistral Small 3.2 specifications
Llama 4 MaverickMistral Small 3.2
ProviderMetaMistral AI
Noometry Index30.931.2
Released2025-04-052025-06-20
WeightsOpenOpen
Context window128K256K
Max output4K16K
Input $ / M tokens$0.19$0.0938
Output $ / M tokens$0.65$0.25
Results tracked546

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Not comparable

Llama 4 Maverick: 26.6 (#324), Mistral Small 3.2: —

Coding benchmarks
BenchmarkLlama 4 MaverickMistral Small 3.2
SWE-bench Verified (bash only)21%—
Aider Polyglot15.6%—
SciCode33.1%—
WeirdML24.5%—
BigCodeBench Instruct49.7%—
LMArena Coding1302—
BigCodeBench Complete61.4%—
ALE-Bench172.97—

Agentic & Tool Use Not comparable

Llama 4 Maverick: 28.2 (#91), Mistral Small 3.2: —

Agentic & Tool Use benchmarks
BenchmarkLlama 4 MaverickMistral Small 3.2
Berkeley Function Calling Leaderboard37.3%—

Reasoning Mistral Small 3.2 leads

Llama 4 Maverick: 10.1 (#342), Mistral Small 3.2: 18.1 (#287)

Reasoning benchmarks
BenchmarkLlama 4 MaverickMistral Small 3.2
Kagi LLM Benchmark55.9%40.4%
Epoch Capabilities Index132.2131.74
ARC-AGI-20%—
SimpleBench27.7%—
NYT Connections (extended)8%—
ARC-AGI-14.4%—
CritPt0%—
Chess Puzzles—1%
EnigmaEval0.6%—
LMArena Hard Prompts1281—
DTBench61.9%—
LMCA15.9%—
ForecastBench57.5—

Math Too close to call

Llama 4 Maverick: 26.0 (#262), Mistral Small 3.2: 26.3 (#260)

Math benchmarks
BenchmarkLlama 4 MaverickMistral Small 3.2
OTIS Mock AIME 2024-202520.6%30.3%
Omni-MATH42.2%—
LMArena Math1299—
MATH Level 573%—
FrontierMath (Feb 2025 set)0.7%—

Knowledge Llama 4 Maverick leads

Llama 4 Maverick: 33.4 (#204), Mistral Small 3.2: 26.7 (#256)

Knowledge benchmarks
BenchmarkLlama 4 MaverickMistral Small 3.2
GPQA Diamond67%49.1%
Humanity's Last Exam5.7%—
MMLU-Pro81%—
Confabulations22.6%—
Vectara Hallucination Rate8.2%—
GPQA (HELM)65%—
LMArena Expert1259—

Multimodal Not comparable

Llama 4 Maverick: 31.6 (#105), Mistral Small 3.2: —

Multimodal benchmarks
BenchmarkLlama 4 MaverickMistral Small 3.2
LMArena Vision1142—
GeoBench52%—
SpatialViz-Bench31.8%—

Multilingual Not comparable

Llama 4 Maverick: 42.2 (#195), Mistral Small 3.2: —

Multilingual benchmarks
BenchmarkLlama 4 MaverickMistral Small 3.2
LMArena Non-English1269—
LMArena Chinese1277—
LMArena French1259—
LMArena German1291—
LMArena Japanese1207—
LMArena Korean1203—
LMArena Russian1286—
LMArena Spanish1293—

Instruction Following Not comparable

Llama 4 Maverick: 71.7 (#146), Mistral Small 3.2: —

Instruction Following benchmarks
BenchmarkLlama 4 MaverickMistral Small 3.2
IFEval90.8%—
LMArena Instruction Following1267—

Long Context Not comparable

Llama 4 Maverick: 31.4 (#279), Mistral Small 3.2: —

Long Context benchmarks
BenchmarkLlama 4 MaverickMistral Small 3.2
Fiction.LiveBench46.2%—
LMArena Longer Query1280—

Writing & Preference Mistral Small 3.2 leads

Llama 4 Maverick: 38.8 (#252), Mistral Small 3.2: 45.0 (#224)

Writing & Preference benchmarks
BenchmarkLlama 4 MaverickMistral Small 3.2
EQ-Bench Creative Writing8601255
LMArena Text1287—
LMArena Creative Writing1267—
Short-Story Creative Writing62%—
WildBench80%—
LMArena Multi-Turn1289—

Frequently asked questions

Is Llama 4 Maverick better than Mistral Small 3.2?

Llama 4 Maverick and Mistral Small 3.2 score almost the same on the Noometry Index (30.9 vs 31.2), so choose on price, context window or the category you care about most.

Which is cheaper, Llama 4 Maverick or Mistral Small 3.2?

Mistral Small 3.2 is cheaper. It lists at $0.0938 per million input tokens and $0.25 per million output tokens; Llama 4 Maverick lists at $0.19 and $0.65.

Which has the bigger context window?

Mistral Small 3.2 does, with 256K tokens against 128K.

How many benchmarks do Llama 4 Maverick and Mistral Small 3.2 share?

5 benchmarks have published results for both models. Llama 4 Maverick has 54 scored results on Noometry and Mistral Small 3.2 has 6.

Related comparisons

Go deeper