Model comparison

Llama 3.2 3B vs Phi-4 Mini

Phi-4 Mini is the stronger model overall, scoring 30.9 to 28.9 on the Noometry Index.

Last verified . 0 shared benchmarks.

Llama 3.2 3B Meta

28.9

Rank #321 Confirmed

Phi-4 Mini Microsoft

30.9

Rank #283 Reported

Summary

  • The widest gap is in knowledge, where Llama 3.2 3B leads 29.7 to 25.3.
  • Llama 3.2 3B is cheaper at $0.05 / $0.33 per million input/output tokens, against $0.075 / $0.30 for Phi-4 Mini.
  • Llama 3.2 3B accepts more context: 131K tokens versus 128K.

Side by side

Llama 3.2 3B and Phi-4 Mini specifications
Llama 3.2 3BPhi-4 Mini
ProviderMetaMicrosoft
Noometry Index28.930.9
Released2024-09-242024-12-11
WeightsOpenOpen
Context window131K128K
Max output118K4K
Input $ / M tokens$0.05$0.075
Output $ / M tokens$0.33$0.30
Results tracked183

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

Llama 3.2 3B: 27.6 (#319), Phi-4 Mini: 28.1 (#317)

Coding benchmarks
BenchmarkLlama 3.2 3BPhi-4 Mini
SciCode—10.8%
BigCodeBench Instruct23.4%—
LMArena Coding1098—
BigCodeBench Complete28.3%—

Agentic & Tool Use Not comparable

Llama 3.2 3B: 20.1 (#143), Phi-4 Mini: —

Agentic & Tool Use benchmarks
BenchmarkLlama 3.2 3BPhi-4 Mini
Berkeley Function Calling Leaderboard21.9%—
BALROG10.1%—

Reasoning Phi-4 Mini leads

Llama 3.2 3B: 21.0 (#228), Phi-4 Mini: 22.4 (#195)

Reasoning benchmarks
BenchmarkLlama 3.2 3BPhi-4 Mini
CritPt—0%
LMArena Hard Prompts1095—

Math Not comparable

Llama 3.2 3B: 32.4 (#214), Phi-4 Mini: —

Math benchmarks
BenchmarkLlama 3.2 3BPhi-4 Mini
LMArena Math1126—

Knowledge Llama 3.2 3B leads

Llama 3.2 3B: 29.7 (#235), Phi-4 Mini: 25.3 (#262)

Knowledge benchmarks
BenchmarkLlama 3.2 3BPhi-4 Mini
Vectara Hallucination Rate—23.5%
LMArena Expert1090—

Multilingual Not comparable

Llama 3.2 3B: 26.2 (#281), Phi-4 Mini: —

Multilingual benchmarks
BenchmarkLlama 3.2 3BPhi-4 Mini
LMArena Non-English1019—
LMArena Chinese1017—
LMArena German1056—
LMArena Russian949—

Instruction Following Not comparable

Llama 3.2 3B: 56.0 (#275), Phi-4 Mini: —

Instruction Following benchmarks
BenchmarkLlama 3.2 3BPhi-4 Mini
LMArena Instruction Following1089—

Long Context Not comparable

Llama 3.2 3B: 33.4 (#261), Phi-4 Mini: —

Long Context benchmarks
BenchmarkLlama 3.2 3BPhi-4 Mini
LMArena Longer Query1100—

Writing & Preference Not comparable

Llama 3.2 3B: 24.7 (#307), Phi-4 Mini: —

Writing & Preference benchmarks
BenchmarkLlama 3.2 3BPhi-4 Mini
LMArena Text1110—
LMArena Creative Writing1094—
EQ-Bench Creative Writing595—
LMArena Multi-Turn1105—

Frequently asked questions

Is Llama 3.2 3B better than Phi-4 Mini?

Phi-4 Mini is the stronger model overall, scoring 30.9 to 28.9 on the Noometry Index.

Which is cheaper, Llama 3.2 3B or Phi-4 Mini?

Llama 3.2 3B is cheaper. It lists at $0.05 per million input tokens and $0.33 per million output tokens; Phi-4 Mini lists at $0.075 and $0.30.

Is Llama 3.2 3B or Phi-4 Mini better for coding?

They score almost the same on coding (27.6 vs 28.1); test both on your own repository before choosing.

Which has the bigger context window?

Llama 3.2 3B does, with 131K tokens against 128K.

How many benchmarks do Llama 3.2 3B and Phi-4 Mini share?

0 benchmarks have published results for both models. Llama 3.2 3B has 18 scored results on Noometry and Phi-4 Mini has 3.

Related comparisons

Go deeper