Model comparison

Mistral Small vs Phi 3 Medium 4k Instruct

Mistral Small and Phi 3 Medium 4k Instruct score almost the same on the Noometry Index (33.4 vs 33.0), so choose on price, context window or the category you care about most.

Last verified . 17 shared benchmarks.

Mistral Small Mistral AI

33.4

Rank #243 Confirmed

Phi 3 Medium 4k Instruct Microsoft

33.0

Rank #250 Confirmed

Summary

  • They share 17 benchmarks with published results for both. Mistral Small scores higher in 6 categories and Phi 3 Medium 4k Instruct in 2 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Mistral Small leads 52.5 to 34.1.

Side by side

Mistral Small and Phi 3 Medium 4k Instruct specifications
Mistral SmallPhi 3 Medium 4k Instruct
ProviderMistral AIMicrosoft
Noometry Index33.433.0
Released2024-02-26—
WeightsOpenOpen
Context window262K—
Max output256K—
Input $ / M tokens$0.15—
Output $ / M tokens$0.60—
Results tracked3917

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Mistral Small leads

Mistral Small: 34.0 (#247), Phi 3 Medium 4k Instruct: 32.9 (#268)

Coding benchmarks
BenchmarkMistral SmallPhi 3 Medium 4k Instruct
LMArena Coding13621130
SciCode26.5%—
BigCodeBench Instruct36.1%—
LiveBench Coding36.2%—
BigCodeBench Complete46.6%—
ALE-Bench497.62—

Agentic & Tool Use Not comparable

Mistral Small: 28.1 (#93), Phi 3 Medium 4k Instruct: —

Agentic & Tool Use benchmarks
BenchmarkMistral SmallPhi 3 Medium 4k Instruct
Berkeley Function Calling Leaderboard37.1%—

Reasoning Phi 3 Medium 4k Instruct leads

Mistral Small: 19.8 (#250), Phi 3 Medium 4k Instruct: 21.7 (#214)

Reasoning benchmarks
BenchmarkMistral SmallPhi 3 Medium 4k Instruct
LMArena Hard Prompts13351127
Kagi LLM Benchmark37.8%—
CritPt0%—
LiveBench Reasoning44.8%—
DTBench70.9%—
LiveBench Data Analysis53.7%—
LMCA20.6%—
LiveBench44%—

Math Phi 3 Medium 4k Instruct leads

Mistral Small: 16.4 (#293), Phi 3 Medium 4k Instruct: 33.4 (#202)

Math benchmarks
BenchmarkMistral SmallPhi 3 Medium 4k Instruct
LMArena Math13411173
OTIS Mock AIME 2024-20255.8%—
LiveBench Math39.9%—
MATH Level 546.8%—

Knowledge Too close to call

Mistral Small: 31.0 (#222), Phi 3 Medium 4k Instruct: 30.2 (#229)

Knowledge benchmarks
BenchmarkMistral SmallPhi 3 Medium 4k Instruct
LMArena Expert12911108
GPQA Diamond47.5%—
Vectara Hallucination Rate5.1%—
MMLU68.7%—

Multimodal Not comparable

Mistral Small: 33.5 (#96), Phi 3 Medium 4k Instruct: —

Multimodal benchmarks
BenchmarkMistral SmallPhi 3 Medium 4k Instruct
LMArena Vision1142—

Multilingual Mistral Small leads

Mistral Small: 45.5 (#169), Phi 3 Medium 4k Instruct: 30.6 (#263)

Multilingual benchmarks
BenchmarkMistral SmallPhi 3 Medium 4k Instruct
LMArena Non-English13151094
LMArena Chinese13401108
LMArena French13371110
LMArena German13401101
LMArena Japanese12751042
LMArena Korean1259954
LMArena Russian13241145
LMArena Spanish13461095

Instruction Following Mistral Small leads

Mistral Small: 66.4 (#209), Phi 3 Medium 4k Instruct: 57.6 (#268)

Instruction Following benchmarks
BenchmarkMistral SmallPhi 3 Medium 4k Instruct
LMArena Instruction Following13101113
LiveBench Instruction Following63.7%—

Long Context Mistral Small leads

Mistral Small: 40.4 (#156), Phi 3 Medium 4k Instruct: 34.0 (#255)

Long Context benchmarks
BenchmarkMistral SmallPhi 3 Medium 4k Instruct
LMArena Longer Query13271120

Writing & Preference Mistral Small leads

Mistral Small: 52.5 (#171), Phi 3 Medium 4k Instruct: 34.1 (#272)

Writing & Preference benchmarks
BenchmarkMistral SmallPhi 3 Medium 4k Instruct
LMArena Text13381138
LMArena Creative Writing13051107
LMArena Multi-Turn13441088
LiveBench Language30.5%—

Frequently asked questions

Is Mistral Small better than Phi 3 Medium 4k Instruct?

Mistral Small and Phi 3 Medium 4k Instruct score almost the same on the Noometry Index (33.4 vs 33.0), so choose on price, context window or the category you care about most.

Is Mistral Small or Phi 3 Medium 4k Instruct better for coding?

Mistral Small scores higher on coding benchmarks: 34.0 versus 32.9 in the Noometry coding category.

How many benchmarks do Mistral Small and Phi 3 Medium 4k Instruct share?

17 benchmarks have published results for both models. Mistral Small has 39 scored results on Noometry and Phi 3 Medium 4k Instruct has 17.

Related comparisons

Go deeper