Model comparison
Llama 3.2 3B vs Phi-4 Mini
Phi-4 Mini is the stronger model overall, scoring 30.9 to 28.9 on the Noometry Index.
Last verified . 0 shared benchmarks.
Summary
- The widest gap is in knowledge, where Llama 3.2 3B leads 29.7 to 25.3.
- Llama 3.2 3B is cheaper at $0.05 / $0.33 per million input/output tokens, against $0.075 / $0.30 for Phi-4 Mini.
- Llama 3.2 3B accepts more context: 131K tokens versus 128K.
Side by side
| Llama 3.2 3B | Phi-4 Mini | |
|---|---|---|
| Provider | Meta | Microsoft |
| Noometry Index | 28.9 | 30.9 |
| Released | 2024-09-24 | 2024-12-11 |
| Weights | Open | Open |
| Context window | 131K | 128K |
| Max output | 118K | 4K |
| Input $ / M tokens | $0.05 | $0.075 |
| Output $ / M tokens | $0.33 | $0.30 |
| Results tracked | 18 | 3 |
Sponsored placements are available on pages like this one. Advertise on Noometry
Category by category
Coding Too close to call
Llama 3.2 3B: 27.6 (#319), Phi-4 Mini: 28.1 (#317)
| Benchmark | Llama 3.2 3B | Phi-4 Mini |
|---|---|---|
| SciCode | — | 10.8% |
| BigCodeBench Instruct | 23.4% | — |
| LMArena Coding | 1098 | — |
| BigCodeBench Complete | 28.3% | — |
Agentic & Tool Use Not comparable
Llama 3.2 3B: 20.1 (#143), Phi-4 Mini: —
| Benchmark | Llama 3.2 3B | Phi-4 Mini |
|---|---|---|
| Berkeley Function Calling Leaderboard | 21.9% | — |
| BALROG | 10.1% | — |
Reasoning Phi-4 Mini leads
Llama 3.2 3B: 21.0 (#228), Phi-4 Mini: 22.4 (#195)
| Benchmark | Llama 3.2 3B | Phi-4 Mini |
|---|---|---|
| CritPt | — | 0% |
| LMArena Hard Prompts | 1095 | — |
Math Not comparable
Llama 3.2 3B: 32.4 (#214), Phi-4 Mini: —
| Benchmark | Llama 3.2 3B | Phi-4 Mini |
|---|---|---|
| LMArena Math | 1126 | — |
Knowledge Llama 3.2 3B leads
Llama 3.2 3B: 29.7 (#235), Phi-4 Mini: 25.3 (#262)
| Benchmark | Llama 3.2 3B | Phi-4 Mini |
|---|---|---|
| Vectara Hallucination Rate | — | 23.5% |
| LMArena Expert | 1090 | — |
Multilingual Not comparable
Llama 3.2 3B: 26.2 (#281), Phi-4 Mini: —
| Benchmark | Llama 3.2 3B | Phi-4 Mini |
|---|---|---|
| LMArena Non-English | 1019 | — |
| LMArena Chinese | 1017 | — |
| LMArena German | 1056 | — |
| LMArena Russian | 949 | — |
Instruction Following Not comparable
Llama 3.2 3B: 56.0 (#275), Phi-4 Mini: —
| Benchmark | Llama 3.2 3B | Phi-4 Mini |
|---|---|---|
| LMArena Instruction Following | 1089 | — |
Long Context Not comparable
Llama 3.2 3B: 33.4 (#261), Phi-4 Mini: —
| Benchmark | Llama 3.2 3B | Phi-4 Mini |
|---|---|---|
| LMArena Longer Query | 1100 | — |
Writing & Preference Not comparable
Llama 3.2 3B: 24.7 (#307), Phi-4 Mini: —
| Benchmark | Llama 3.2 3B | Phi-4 Mini |
|---|---|---|
| LMArena Text | 1110 | — |
| LMArena Creative Writing | 1094 | — |
| EQ-Bench Creative Writing | 595 | — |
| LMArena Multi-Turn | 1105 | — |
Frequently asked questions
Is Llama 3.2 3B better than Phi-4 Mini?
Phi-4 Mini is the stronger model overall, scoring 30.9 to 28.9 on the Noometry Index.
Which is cheaper, Llama 3.2 3B or Phi-4 Mini?
Llama 3.2 3B is cheaper. It lists at $0.05 per million input tokens and $0.33 per million output tokens; Phi-4 Mini lists at $0.075 and $0.30.
Is Llama 3.2 3B or Phi-4 Mini better for coding?
They score almost the same on coding (27.6 vs 28.1); test both on your own repository before choosing.
Which has the bigger context window?
Llama 3.2 3B does, with 131K tokens against 128K.
How many benchmarks do Llama 3.2 3B and Phi-4 Mini share?
0 benchmarks have published results for both models. Llama 3.2 3B has 18 scored results on Noometry and Phi-4 Mini has 3.