Model comparison

Mistral Large vs Phi-4

Mistral Large and Phi-4 score almost the same on the Noometry Index (31.9 vs 31.2), so choose on price, context window or the category you care about most.

Last verified . 35 shared benchmarks.

Mistral Large Mistral AI

31.9

Rank #263 Confirmed

Phi-4 Microsoft

31.2

Rank #279 Confirmed

Summary

  • They share 35 benchmarks with published results for both. Mistral Large scores higher in 5 categories and Phi-4 in 4 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in instruction following, where Mistral Large leads 67.9 to 60.4.
  • The biggest single-benchmark swing is BigCodeBench Complete: 38.3% for Mistral Large and 55.4% for Phi-4.
  • Phi-4 is cheaper at $0.07 / $0.14 per million input/output tokens, against $2 / $6 for Mistral Large.
  • Mistral Large accepts more context: 131K tokens versus 128K.

Side by side

Mistral Large and Phi-4 specifications
Mistral LargePhi-4
ProviderMistral AIMicrosoft
Noometry Index31.931.2
Released2024-02-262024-12-11
WeightsOpenOpen
Context window131K128K
Max output16K4K
Input $ / M tokens$2$0.07
Output $ / M tokens$6$0.14
Results tracked5137

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

Mistral Large: 34.3 (#240), Phi-4: 34.4 (#239)

Coding benchmarks
BenchmarkMistral LargePhi-4
BigCodeBench Instruct30%45.5%
LiveBench Coding47.1%30.7%
LMArena Coding12771231
BigCodeBench Complete38.3%55.4%
SciCode36.2%—
ALE-Bench264.7—
HumanEval+62.2%—
MBPP+59.5%—

Agentic & Tool Use Mistral Large leads

Mistral Large: 28.6 (#89), Phi-4: 22.8 (#128)

Agentic & Tool Use benchmarks
BenchmarkMistral LargePhi-4
Berkeley Function Calling Leaderboard38.4%28.8%
BALROG—11.6%

Reasoning Phi-4 leads

Mistral Large: 15.8 (#310), Phi-4: 17.7 (#291)

Reasoning benchmarks
BenchmarkMistral LargePhi-4
LiveBench Reasoning43.5%47.8%
LMArena Hard Prompts12571220
LiveBench Data Analysis50.1%45.2%
Epoch Capabilities Index128.52130.42
LiveBench48.4%41.6%
SimpleBench22.5%—
CritPt0%—
Chess Puzzles—1%
DTBench65.1%—
LMCA16.7%—
ForecastBench57.1—

Math Phi-4 leads

Mistral Large: 18.2 (#291), Phi-4: 20.8 (#285)

Math benchmarks
BenchmarkMistral LargePhi-4
OTIS Mock AIME 2024-20258.5%13.8%
LiveBench Math42.5%42%
LMArena Math12621246
MATH Level 550.3%64.9%
Omni-MATH28.1%—
FrontierMath (Feb 2025 set)0.3%—

Knowledge Phi-4 leads

Mistral Large: 30.1 (#230), Phi-4: 32.6 (#209)

Knowledge benchmarks
BenchmarkMistral LargePhi-4
GPQA Diamond51.3%56.1%
Confabulations21.4%29.4%
Vectara Hallucination Rate4.5%3.7%
LMArena Expert12321203
MMLU80%84.8%
MMLU-Pro59.9%—
GPQA (HELM)43.5%—

Multilingual Mistral Large leads

Mistral Large: 40.0 (#219), Phi-4: 37.2 (#237)

Multilingual benchmarks
BenchmarkMistral LargePhi-4
LMArena Non-English12371197
LMArena Chinese12401212
LMArena French13251224
LMArena German12541222
LMArena Japanese11881158
LMArena Korean12021151
LMArena Russian12571209
LMArena Spanish12681234

Instruction Following Mistral Large leads

Mistral Large: 67.9 (#191), Phi-4: 60.4 (#251)

Instruction Following benchmarks
BenchmarkMistral LargePhi-4
LiveBench Instruction Following67.9%58.4%
LMArena Instruction Following12491201
IFEval87.7%—

Long Context Mistral Large leads

Mistral Large: 38.3 (#199), Phi-4: 36.9 (#226)

Long Context benchmarks
BenchmarkMistral LargePhi-4
LMArena Longer Query12611217

Writing & Preference Too close to call

Mistral Large: 40.7 (#242), Phi-4: 40.5 (#244)

Writing & Preference benchmarks
BenchmarkMistral LargePhi-4
LMArena Text12661217
LMArena Creative Writing12431182
Short-Story Creative Writing69%62.6%
LMArena Multi-Turn12601206
LiveBench Language39.4%25.6%
EQ-Bench Creative Writing985—
WildBench80.1%—

Frequently asked questions

Is Mistral Large better than Phi-4?

Mistral Large and Phi-4 score almost the same on the Noometry Index (31.9 vs 31.2), so choose on price, context window or the category you care about most.

Which is cheaper, Mistral Large or Phi-4?

Phi-4 is cheaper. It lists at $0.07 per million input tokens and $0.14 per million output tokens; Mistral Large lists at $2 and $6.

Is Mistral Large or Phi-4 better for coding?

They score almost the same on coding (34.3 vs 34.4); test both on your own repository before choosing.

Which has the bigger context window?

Mistral Large does, with 131K tokens against 128K.

How many benchmarks do Mistral Large and Phi-4 share?

35 benchmarks have published results for both models. Mistral Large has 51 scored results on Noometry and Phi-4 has 37.

Related comparisons

Go deeper