Model comparison

Llama 4 Maverick vs Mistral Small 3

Llama 4 Maverick and Mistral Small 3 score almost the same on the Noometry Index (30.9 vs 31.2), so choose on price, context window or the category you care about most.

Last verified . 23 shared benchmarks.

Llama 4 Maverick Meta

30.9

Rank #282 Confirmed

Mistral Small 3 Mistral AI

31.2

Rank #278 Confirmed

Summary

  • They share 23 benchmarks with published results for both. Llama 4 Maverick scores higher in 5 categories and Mistral Small 3 in 3 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in coding, where Mistral Small 3 leads 36.5 to 26.6.
  • The biggest single-benchmark swing is GPQA Diamond: 67% for Llama 4 Maverick and 47.3% for Mistral Small 3.
  • Mistral Small 3 is cheaper at $0.05 / $0.08 per million input/output tokens, against $0.19 / $0.65 for Llama 4 Maverick.
  • Llama 4 Maverick accepts more context: 128K tokens versus 33K.

Side by side

Llama 4 Maverick and Mistral Small 3 specifications
Llama 4 MaverickMistral Small 3
ProviderMetaMistral AI
Noometry Index30.931.2
Released2025-04-052025-01-30
WeightsOpenOpen
Context window128K33K
Max output4K16K
Input $ / M tokens$0.19$0.05
Output $ / M tokens$0.65$0.08
Results tracked5424

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Mistral Small 3 leads

Llama 4 Maverick: 26.6 (#324), Mistral Small 3: 36.5 (#207)

Coding benchmarks
BenchmarkLlama 4 MaverickMistral Small 3
BigCodeBench Instruct49.7%45.3%
LMArena Coding13021246
BigCodeBench Complete61.4%50.4%
SWE-bench Verified (bash only)21%—
Aider Polyglot15.6%—
SciCode33.1%—
WeirdML24.5%—
ALE-Bench172.97—

Agentic & Tool Use Not comparable

Llama 4 Maverick: 28.2 (#91), Mistral Small 3: —

Agentic & Tool Use benchmarks
BenchmarkLlama 4 MaverickMistral Small 3
Berkeley Function Calling Leaderboard37.3%—

Reasoning Mistral Small 3 leads

Llama 4 Maverick: 10.1 (#342), Mistral Small 3: 18.9 (#273)

Reasoning benchmarks
BenchmarkLlama 4 MaverickMistral Small 3
LMArena Hard Prompts12811233
Epoch Capabilities Index132.2127.07
ARC-AGI-20%—
SimpleBench27.7%—
Kagi LLM Benchmark55.9%—
NYT Connections (extended)8%—
ARC-AGI-14.4%—
CritPt0%—
Chess Puzzles—0%
EnigmaEval0.6%—
DTBench61.9%—
LMCA15.9%—
ForecastBench57.5—

Math Llama 4 Maverick leads

Llama 4 Maverick: 26.0 (#262), Mistral Small 3: 16.3 (#295)

Math benchmarks
BenchmarkLlama 4 MaverickMistral Small 3
OTIS Mock AIME 2024-202520.6%6.7%
LMArena Math12991240
Omni-MATH42.2%—
MATH Level 573%—
FrontierMath (Feb 2025 set)0.7%—

Knowledge Llama 4 Maverick leads

Llama 4 Maverick: 33.4 (#204), Mistral Small 3: 25.1 (#263)

Knowledge benchmarks
BenchmarkLlama 4 MaverickMistral Small 3
GPQA Diamond67%47.3%
Confabulations22.6%25.2%
LMArena Expert12591202
Humanity's Last Exam5.7%—
MMLU-Pro81%—
Vectara Hallucination Rate8.2%—
GPQA (HELM)65%—

Multimodal Not comparable

Llama 4 Maverick: 31.6 (#105), Mistral Small 3: —

Multimodal benchmarks
BenchmarkLlama 4 MaverickMistral Small 3
LMArena Vision1142—
GeoBench52%—
SpatialViz-Bench31.8%—

Multilingual Llama 4 Maverick leads

Llama 4 Maverick: 42.2 (#195), Mistral Small 3: 37.3 (#236)

Multilingual benchmarks
BenchmarkLlama 4 MaverickMistral Small 3
LMArena Non-English12691198
LMArena Chinese12771204
LMArena French12591203
LMArena German12911211
LMArena Japanese12071111
LMArena Korean12031188
LMArena Russian12861216
LMArena Spanish1293—

Instruction Following Llama 4 Maverick leads

Llama 4 Maverick: 71.7 (#146), Mistral Small 3: 63.7 (#229)

Instruction Following benchmarks
BenchmarkLlama 4 MaverickMistral Small 3
LMArena Instruction Following12671214
IFEval90.8%—

Long Context Mistral Small 3 leads

Llama 4 Maverick: 31.4 (#279), Mistral Small 3: 37.8 (#211)

Long Context benchmarks
BenchmarkLlama 4 MaverickMistral Small 3
LMArena Longer Query12801246
Fiction.LiveBench46.2%—

Writing & Preference Llama 4 Maverick leads

Llama 4 Maverick: 38.8 (#252), Mistral Small 3: 32.2 (#280)

Writing & Preference benchmarks
BenchmarkLlama 4 MaverickMistral Small 3
LMArena Text12871234
LMArena Creative Writing12671195
EQ-Bench Creative Writing860707
LMArena Multi-Turn12891217
Short-Story Creative Writing62%—
WildBench80%—

Frequently asked questions

Is Llama 4 Maverick better than Mistral Small 3?

Llama 4 Maverick and Mistral Small 3 score almost the same on the Noometry Index (30.9 vs 31.2), so choose on price, context window or the category you care about most.

Which is cheaper, Llama 4 Maverick or Mistral Small 3?

Mistral Small 3 is cheaper. It lists at $0.05 per million input tokens and $0.08 per million output tokens; Llama 4 Maverick lists at $0.19 and $0.65.

Is Llama 4 Maverick or Mistral Small 3 better for coding?

Mistral Small 3 scores higher on coding benchmarks: 36.5 versus 26.6 in the Noometry coding category.

Which has the bigger context window?

Llama 4 Maverick does, with 128K tokens against 33K.

How many benchmarks do Llama 4 Maverick and Mistral Small 3 share?

23 benchmarks have published results for both models. Llama 4 Maverick has 54 scored results on Noometry and Mistral Small 3 has 24.

Related comparisons

Go deeper