Model comparison

Llama 3.2 1B vs Phi-4 Mini

Phi-4 Mini is the stronger model overall, scoring 30.9 to 20.1 on the Noometry Index. Llama 3.2 1B costs 1.9× less per token, which makes it the better buy when Phi-4 Mini's lead doesn't matter for your workload.

Last verified . 0 shared benchmarks.

Llama 3.2 1B Meta

20.1

Rank #354 Confirmed

Phi-4 Mini Microsoft

30.9

Rank #283 Reported

Summary

  • The widest gap is in knowledge, where Phi-4 Mini leads 25.3 to 7.2.
  • Llama 3.2 1B is cheaper at $0.027 / $0.20 per million input/output tokens, against $0.075 / $0.30 for Phi-4 Mini.
  • Phi-4 Mini accepts more context: 128K tokens versus 60K.

Side by side

Llama 3.2 1B and Phi-4 Mini specifications
Llama 3.2 1BPhi-4 Mini
ProviderMetaMicrosoft
Noometry Index20.130.9
Released2024-09-242024-12-11
WeightsOpenOpen
Context window60K128K
Max output54K4K
Input $ / M tokens$0.027$0.075
Output $ / M tokens$0.20$0.30
Results tracked223

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Phi-4 Mini leads

Llama 3.2 1B: 21.1 (#338), Phi-4 Mini: 28.1 (#317)

Coding benchmarks
BenchmarkLlama 3.2 1BPhi-4 Mini
SciCode—10.8%
BigCodeBench Instruct8.2%—
LMArena Coding1070—
BigCodeBench Complete11.3%—

Agentic & Tool Use Not comparable

Llama 3.2 1B: 14.6 (#150), Phi-4 Mini: —

Agentic & Tool Use benchmarks
BenchmarkLlama 3.2 1BPhi-4 Mini
Berkeley Function Calling Leaderboard10.8%—
BALROG6.6%—

Reasoning Phi-4 Mini leads

Llama 3.2 1B: 16.2 (#308), Phi-4 Mini: 22.4 (#195)

Reasoning benchmarks
BenchmarkLlama 3.2 1BPhi-4 Mini
CritPt—0%
Chess Puzzles0%—
LMArena Hard Prompts1044—
Epoch Capabilities Index101.99—

Math Not comparable

Llama 3.2 1B: 10.4 (#313), Phi-4 Mini: —

Math benchmarks
BenchmarkLlama 3.2 1BPhi-4 Mini
OTIS Mock AIME 2024-20250.6%—
LMArena Math1086—

Knowledge Phi-4 Mini leads

Llama 3.2 1B: 7.2 (#312), Phi-4 Mini: 25.3 (#262)

Knowledge benchmarks
BenchmarkLlama 3.2 1BPhi-4 Mini
GPQA Diamond23.9%—
Vectara Hallucination Rate—23.5%
LMArena Expert1007—

Multilingual Not comparable

Llama 3.2 1B: 23.8 (#292), Phi-4 Mini: —

Multilingual benchmarks
BenchmarkLlama 3.2 1BPhi-4 Mini
LMArena Non-English973—
LMArena Chinese959—
LMArena German1014—
LMArena Russian941—

Instruction Following Not comparable

Llama 3.2 1B: 52.4 (#290), Phi-4 Mini: —

Instruction Following benchmarks
BenchmarkLlama 3.2 1BPhi-4 Mini
LMArena Instruction Following1031—

Long Context Not comparable

Llama 3.2 1B: 31.9 (#274), Phi-4 Mini: —

Long Context benchmarks
BenchmarkLlama 3.2 1BPhi-4 Mini
LMArena Longer Query1050—

Writing & Preference Not comparable

Llama 3.2 1B: 21.3 (#310), Phi-4 Mini: —

Writing & Preference benchmarks
BenchmarkLlama 3.2 1BPhi-4 Mini
LMArena Text1055—
LMArena Creative Writing1033—
EQ-Bench Creative Writing200—
LMArena Multi-Turn1030—

Frequently asked questions

Is Llama 3.2 1B better than Phi-4 Mini?

Phi-4 Mini is the stronger model overall, scoring 30.9 to 20.1 on the Noometry Index. Llama 3.2 1B costs 1.9× less per token, which makes it the better buy when Phi-4 Mini's lead doesn't matter for your workload.

Which is cheaper, Llama 3.2 1B or Phi-4 Mini?

Llama 3.2 1B is cheaper. It lists at $0.027 per million input tokens and $0.20 per million output tokens; Phi-4 Mini lists at $0.075 and $0.30.

Is Llama 3.2 1B or Phi-4 Mini better for coding?

Phi-4 Mini scores higher on coding benchmarks: 28.1 versus 21.1 in the Noometry coding category.

Which has the bigger context window?

Phi-4 Mini does, with 128K tokens against 60K.

How many benchmarks do Llama 3.2 1B and Phi-4 Mini share?

0 benchmarks have published results for both models. Llama 3.2 1B has 22 scored results on Noometry and Phi-4 Mini has 3.

Related comparisons

Go deeper