Model comparison

Llama 3-8B vs Ministral 3B

Llama 3-8B and Ministral 3B score almost the same on the Noometry Index (25.5 vs 26.2), so choose on price, context window or the category you care about most.

Last verified . 4 shared benchmarks.

Llama 3-8B Meta

25.5

Rank #344 Confirmed

Ministral 3B Mistral AI

26.2

Rank #338 Confirmed

Summary

  • They share 4 benchmarks with published results for both. Llama 3-8B scores higher in 0 categories and Ministral 3B in 3 categories; 3 gaps are clear of the uncertainty.
  • The widest gap is in math, where Ministral 3B leads 26.6 to 8.8.
  • The biggest single-benchmark swing is MATH Level 5: 6.1% for Llama 3-8B and 14.4% for Ministral 3B.

Side by side

Llama 3-8B and Ministral 3B specifications
Llama 3-8BMinistral 3B
ProviderMetaMistral AI
Noometry Index25.526.2
Released2024-04-182024-10-01
WeightsOpenOpen
Context window—131K
Max output—262K
Input $ / M tokens—$0.10
Output $ / M tokens—$0.10
Results tracked346

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Not comparable

Llama 3-8B: 31.0 (#289), Ministral 3B: —

Coding benchmarks
BenchmarkLlama 3-8BMinistral 3B
BigCodeBench Instruct31.9%—
LMArena Coding1152—
BigCodeBench Complete36.9%—
HumanEval+56.7%—
MBPP+54.8%—

Reasoning Ministral 3B leads

Llama 3-8B: 14.3 (#326), Ministral 3B: 18.4 (#282)

Reasoning benchmarks
BenchmarkLlama 3-8BMinistral 3B
DTBench43.9%51.7%
Epoch Capabilities Index116.45118.1
Chess Puzzles0%—
LMArena Hard Prompts1133—
LMCA—5.5%
Adversarial NLI57.3%—
ForecastBench58.6—
WinoGrande75.7%—

Math Ministral 3B leads

Llama 3-8B: 8.8 (#323), Ministral 3B: 26.6 (#258)

Math benchmarks
BenchmarkLlama 3-8BMinistral 3B
MATH Level 56.1%14.4%
OTIS Mock AIME 2024-20251.9%—
LMArena Math1151—

Knowledge Ministral 3B leads

Llama 3-8B: 7.8 (#308), Ministral 3B: 10.4 (#302)

Knowledge benchmarks
BenchmarkLlama 3-8BMinistral 3B
GPQA Diamond26.1%25.3%
Vectara Hallucination Rate—7.3%
LMArena Expert1113—
ARC (AI2) Challenge82.8%—
MMLU68.8%—
OpenBookQA82.6%—
TriviaQA67.7%—

Multilingual Not comparable

Llama 3-8B: 30.8 (#261), Ministral 3B: —

Multilingual benchmarks
BenchmarkLlama 3-8BMinistral 3B
LMArena Non-English1098—
LMArena Chinese1076—
LMArena French1159—
LMArena German1104—
LMArena Japanese967—
LMArena Korean1004—
LMArena Russian1109—
LMArena Spanish1173—

Instruction Following Not comparable

Llama 3-8B: 58.4 (#260), Ministral 3B: —

Instruction Following benchmarks
BenchmarkLlama 3-8BMinistral 3B
LMArena Instruction Following1127—

Long Context Not comparable

Llama 3-8B: 34.2 (#251), Ministral 3B: —

Long Context benchmarks
BenchmarkLlama 3-8BMinistral 3B
LMArena Longer Query1128—

Writing & Preference Not comparable

Llama 3-8B: 37.5 (#256), Ministral 3B: —

Writing & Preference benchmarks
BenchmarkLlama 3-8BMinistral 3B
LMArena Text1166—
LMArena Creative Writing1150—
LMArena Multi-Turn1152—

Frequently asked questions

Is Llama 3-8B better than Ministral 3B?

Llama 3-8B and Ministral 3B score almost the same on the Noometry Index (25.5 vs 26.2), so choose on price, context window or the category you care about most.

How many benchmarks do Llama 3-8B and Ministral 3B share?

4 benchmarks have published results for both models. Llama 3-8B has 34 scored results on Noometry and Ministral 3B has 6.

Related comparisons

Go deeper