Model comparison

Ministral 8B vs Phi-4 Mini

Phi-4 Mini is the stronger model overall, scoring 30.9 to 28.2 on the Noometry Index.

Last verified . 1 shared benchmarks.

Ministral 8B Mistral AI

28.2

Rank #325 Confirmed

Phi-4 Mini Microsoft

30.9

Rank #283 Reported

Summary

  • They share 1 benchmark with published results for both. Ministral 8B scores higher in 1 category and Phi-4 Mini in 2 categories; 3 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Phi-4 Mini leads 25.3 to 12.6.
  • The biggest single-benchmark swing is Vectara Hallucination Rate: 7.4% for Ministral 8B and 23.5% for Phi-4 Mini.
  • Phi-4 Mini is cheaper at $0.075 / $0.30 per million input/output tokens, against $0.15 / $0.15 for Ministral 8B.
  • Ministral 8B accepts more context: 262K tokens versus 128K.

Side by side

Ministral 8B and Phi-4 Mini specifications
Ministral 8BPhi-4 Mini
ProviderMistral AIMicrosoft
Noometry Index28.230.9
Released2024-10-012024-12-11
WeightsOpenOpen
Context window262K128K
Max output262K4K
Input $ / M tokens$0.15$0.075
Output $ / M tokens$0.15$0.30
Results tracked173

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Ministral 8B leads

Ministral 8B: 35.0 (#230), Phi-4 Mini: 28.1 (#317)

Coding benchmarks
BenchmarkMinistral 8BPhi-4 Mini
SciCode—10.8%
LMArena Coding1202—

Agentic & Tool Use Not comparable

Ministral 8B: 16.4 (#148), Phi-4 Mini: —

Agentic & Tool Use benchmarks
BenchmarkMinistral 8BPhi-4 Mini
Berkeley Function Calling Leaderboard11.1%—

Reasoning Phi-4 Mini leads

Ministral 8B: 18.4 (#281), Phi-4 Mini: 22.4 (#195)

Reasoning benchmarks
BenchmarkMinistral 8BPhi-4 Mini
CritPt—0%
LMArena Hard Prompts1191—
DTBench45.7%—

Math Not comparable

Ministral 8B: 25.7 (#267), Phi-4 Mini: —

Math benchmarks
BenchmarkMinistral 8BPhi-4 Mini
LMArena Math1188—
MATH Level 514.9%—

Knowledge Phi-4 Mini leads

Ministral 8B: 12.6 (#297), Phi-4 Mini: 25.3 (#262)

Knowledge benchmarks
BenchmarkMinistral 8BPhi-4 Mini
Vectara Hallucination Rate7.4%23.5%
GPQA Diamond27.1%—
LMArena Expert1170—

Multilingual Not comparable

Ministral 8B: 35.1 (#247), Phi-4 Mini: —

Multilingual benchmarks
BenchmarkMinistral 8BPhi-4 Mini
LMArena Non-English1165—
LMArena Chinese1193—
LMArena Russian1195—

Instruction Following Not comparable

Ministral 8B: 60.5 (#250), Phi-4 Mini: —

Instruction Following benchmarks
BenchmarkMinistral 8BPhi-4 Mini
LMArena Instruction Following1161—

Long Context Not comparable

Ministral 8B: 36.7 (#227), Phi-4 Mini: —

Long Context benchmarks
BenchmarkMinistral 8BPhi-4 Mini
LMArena Longer Query1212—

Writing & Preference Not comparable

Ministral 8B: 39.6 (#246), Phi-4 Mini: —

Writing & Preference benchmarks
BenchmarkMinistral 8BPhi-4 Mini
LMArena Text1191—
LMArena Creative Writing1175—
LMArena Multi-Turn1166—

Frequently asked questions

Is Ministral 8B better than Phi-4 Mini?

Phi-4 Mini is the stronger model overall, scoring 30.9 to 28.2 on the Noometry Index.

Which is cheaper, Ministral 8B or Phi-4 Mini?

Phi-4 Mini is cheaper. It lists at $0.075 per million input tokens and $0.30 per million output tokens; Ministral 8B lists at $0.15 and $0.15.

Is Ministral 8B or Phi-4 Mini better for coding?

Ministral 8B scores higher on coding benchmarks: 35.0 versus 28.1 in the Noometry coding category.

Which has the bigger context window?

Ministral 8B does, with 262K tokens against 128K.

How many benchmarks do Ministral 8B and Phi-4 Mini share?

1 benchmark has published results for both models. Ministral 8B has 17 scored results on Noometry and Phi-4 Mini has 3.

Related comparisons

Go deeper