Model comparison

Llama 3.2 3B vs Phi-4

Phi-4 is the stronger model overall, scoring 31.2 to 28.9 on the Noometry Index.

Last verified . 17 shared benchmarks.

Llama 3.2 3B Meta

28.9

Rank #321 Confirmed

Phi-4 Microsoft

31.2

Rank #279 Confirmed

Summary

  • They share 17 benchmarks with published results for both. Llama 3.2 3B scores higher in 2 categories and Phi-4 in 7 categories; 9 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Phi-4 leads 40.5 to 24.7.
  • The biggest single-benchmark swing is BigCodeBench Complete: 28.3% for Llama 3.2 3B and 55.4% for Phi-4.
  • Phi-4 is cheaper at $0.07 / $0.14 per million input/output tokens, against $0.05 / $0.33 for Llama 3.2 3B.
  • Llama 3.2 3B accepts more context: 131K tokens versus 128K.

Side by side

Llama 3.2 3B and Phi-4 specifications
Llama 3.2 3BPhi-4
ProviderMetaMicrosoft
Noometry Index28.931.2
Released2024-09-242024-12-11
WeightsOpenOpen
Context window131K128K
Max output118K4K
Input $ / M tokens$0.05$0.07
Output $ / M tokens$0.33$0.14
Results tracked1837

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Phi-4 leads

Llama 3.2 3B: 27.6 (#319), Phi-4: 34.4 (#239)

Coding benchmarks
BenchmarkLlama 3.2 3BPhi-4
BigCodeBench Instruct23.4%45.5%
LMArena Coding10981231
BigCodeBench Complete28.3%55.4%
LiveBench Coding—30.7%

Agentic & Tool Use Phi-4 leads

Llama 3.2 3B: 20.1 (#143), Phi-4: 22.8 (#128)

Agentic & Tool Use benchmarks
BenchmarkLlama 3.2 3BPhi-4
Berkeley Function Calling Leaderboard21.9%28.8%
BALROG10.1%11.6%

Reasoning Llama 3.2 3B leads

Llama 3.2 3B: 21.0 (#228), Phi-4: 17.7 (#291)

Reasoning benchmarks
BenchmarkLlama 3.2 3BPhi-4
LMArena Hard Prompts10951220
Chess Puzzles—1%
LiveBench Reasoning—47.8%
LiveBench Data Analysis—45.2%
Epoch Capabilities Index—130.42
LiveBench—41.6%

Math Llama 3.2 3B leads

Llama 3.2 3B: 32.4 (#214), Phi-4: 20.8 (#285)

Math benchmarks
BenchmarkLlama 3.2 3BPhi-4
LMArena Math11261246
OTIS Mock AIME 2024-2025—13.8%
LiveBench Math—42%
MATH Level 5—64.9%

Knowledge Phi-4 leads

Llama 3.2 3B: 29.7 (#235), Phi-4: 32.6 (#209)

Knowledge benchmarks
BenchmarkLlama 3.2 3BPhi-4
LMArena Expert10901203
GPQA Diamond—56.1%
Confabulations—29.4%
Vectara Hallucination Rate—3.7%
MMLU—84.8%

Multilingual Phi-4 leads

Llama 3.2 3B: 26.2 (#281), Phi-4: 37.2 (#237)

Multilingual benchmarks
BenchmarkLlama 3.2 3BPhi-4
LMArena Non-English10191197
LMArena Chinese10171212
LMArena German10561222
LMArena Russian9491209
LMArena French—1224
LMArena Japanese—1158
LMArena Korean—1151
LMArena Spanish—1234

Instruction Following Phi-4 leads

Llama 3.2 3B: 56.0 (#275), Phi-4: 60.4 (#251)

Instruction Following benchmarks
BenchmarkLlama 3.2 3BPhi-4
LMArena Instruction Following10891201
LiveBench Instruction Following—58.4%

Long Context Phi-4 leads

Llama 3.2 3B: 33.4 (#261), Phi-4: 36.9 (#226)

Long Context benchmarks
BenchmarkLlama 3.2 3BPhi-4
LMArena Longer Query11001217

Writing & Preference Phi-4 leads

Llama 3.2 3B: 24.7 (#307), Phi-4: 40.5 (#244)

Writing & Preference benchmarks
BenchmarkLlama 3.2 3BPhi-4
LMArena Text11101217
LMArena Creative Writing10941182
LMArena Multi-Turn11051206
Short-Story Creative Writing—62.6%
EQ-Bench Creative Writing595—
LiveBench Language—25.6%

Frequently asked questions

Is Llama 3.2 3B better than Phi-4?

Phi-4 is the stronger model overall, scoring 31.2 to 28.9 on the Noometry Index.

Which is cheaper, Llama 3.2 3B or Phi-4?

Phi-4 is cheaper. It lists at $0.07 per million input tokens and $0.14 per million output tokens; Llama 3.2 3B lists at $0.05 and $0.33.

Is Llama 3.2 3B or Phi-4 better for coding?

Phi-4 scores higher on coding benchmarks: 34.4 versus 27.6 in the Noometry coding category.

Which has the bigger context window?

Llama 3.2 3B does, with 131K tokens against 128K.

How many benchmarks do Llama 3.2 3B and Phi-4 share?

17 benchmarks have published results for both models. Llama 3.2 3B has 18 scored results on Noometry and Phi-4 has 37.

Related comparisons

Go deeper