Model comparison

GPT-5 Nano vs Phi-4

GPT-5 Nano is the stronger model overall, scoring 33.5 to 31.2 on the Noometry Index. Phi-4 costs 1.6× less per token, which makes it the better buy when GPT-5 Nano's lead doesn't matter for your workload.

Last verified . 23 shared benchmarks.

GPT-5 Nano OpenAI

33.5

Rank #241 Confirmed

Phi-4 Microsoft

31.2

Rank #279 Confirmed

Summary

  • They share 23 benchmarks with published results for both. GPT-5 Nano scores higher in 5 categories and Phi-4 in 4 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in instruction following, where GPT-5 Nano leads 75.0 to 60.4.
  • The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 81.1% for GPT-5 Nano and 13.8% for Phi-4.
  • Phi-4 is cheaper at $0.07 / $0.14 per million input/output tokens, against $0.05 / $0.40 for GPT-5 Nano.
  • GPT-5 Nano accepts more context: 400K tokens versus 128K.
  • Phi-4 has downloadable open weights; the other is API-only.

Side by side

GPT-5 Nano and Phi-4 specifications
GPT-5 NanoPhi-4
ProviderOpenAIMicrosoft
Noometry Index33.531.2
Released2025-08-072024-12-11
WeightsProprietaryOpen
Context window400K128K
Max output128K4K
Input $ / M tokens$0.05$0.07
Output $ / M tokens$0.40$0.14
Results tracked4937

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

GPT-5 Nano: 33.6 (#254), Phi-4: 34.4 (#239)

Coding benchmarks
BenchmarkGPT-5 NanoPhi-4
LMArena Coding13511231
SWE-bench Verified (bash only)34.8%—
WeirdML38.1%—
BigCodeBench Instruct—45.5%
LiveBench Coding—30.7%
BigCodeBench Complete—55.4%
ALE-Bench718.67—

Agentic & Tool Use GPT-5 Nano leads

GPT-5 Nano: 25.8 (#106), Phi-4: 22.8 (#128)

Agentic & Tool Use benchmarks
BenchmarkGPT-5 NanoPhi-4
Berkeley Function Calling Leaderboard51.5%28.8%
Terminal-Bench21.8%—
BALROG—11.6%

Reasoning Phi-4 leads

GPT-5 Nano: 16.3 (#306), Phi-4: 17.7 (#291)

Reasoning benchmarks
BenchmarkGPT-5 NanoPhi-4
Chess Puzzles27%1%
LMArena Hard Prompts13281220
Epoch Capabilities Index139.38130.42
ARC-AGI-22.6%—
Kagi LLM Benchmark62.2%—
ARC-AGI-120.7%—
LiveBench Reasoning—47.8%
Mystery Game Puzzles9%—
DTBench62.7%—
LiveBench Data Analysis—45.2%
LMCA7.9%—
ForecastBench59.1—
LiveBench—41.6%

Math GPT-5 Nano leads

GPT-5 Nano: 29.4 (#241), Phi-4: 20.8 (#285)

Knowledge GPT-5 Nano leads

GPT-5 Nano: 35.9 (#178), Phi-4: 32.6 (#209)

Knowledge benchmarks
BenchmarkGPT-5 NanoPhi-4
GPQA Diamond69.4%56.1%
Vectara Hallucination Rate10.5%3.7%
LMArena Expert13211203
SimpleQA Verified11.7%—
MMLU-Pro77.8%—
Confabulations—29.4%
GPQA (HELM)67.9%—
MMLU—84.8%

Multimodal Not comparable

GPT-5 Nano: 31.3 (#108), Phi-4: —

Multimodal benchmarks
BenchmarkGPT-5 NanoPhi-4
LMArena Vision1159—
VPCT37.2%—

Multilingual GPT-5 Nano leads

GPT-5 Nano: 45.3 (#172), Phi-4: 37.2 (#237)

Multilingual benchmarks
BenchmarkGPT-5 NanoPhi-4
LMArena Non-English13131197
LMArena Chinese13561212
LMArena German13271222
LMArena Japanese12261158
LMArena Korean12691151
LMArena Russian12961209
LMArena Spanish13601234
LMArena French—1224

Instruction Following GPT-5 Nano leads

GPT-5 Nano: 75.0 (#79), Phi-4: 60.4 (#251)

Instruction Following benchmarks
BenchmarkGPT-5 NanoPhi-4
LMArena Instruction Following13061201
LiveBench Instruction Following—58.4%
IFEval93.2%—

Long Context Phi-4 leads

GPT-5 Nano: 31.3 (#281), Phi-4: 36.9 (#226)

Long Context benchmarks
BenchmarkGPT-5 NanoPhi-4
LMArena Longer Query13121217
Fiction.LiveBench44.4%—

Writing & Preference Phi-4 leads

GPT-5 Nano: 39.1 (#249), Phi-4: 40.5 (#244)

Writing & Preference benchmarks
BenchmarkGPT-5 NanoPhi-4
LMArena Text13201217
LMArena Creative Writing12491182
LMArena Multi-Turn13111206
Short-Story Creative Writing—62.6%
EQ-Bench Creative Writing705—
WildBench80.6%—
LiveBench Language—25.6%

Frequently asked questions

Is GPT-5 Nano better than Phi-4?

GPT-5 Nano is the stronger model overall, scoring 33.5 to 31.2 on the Noometry Index. Phi-4 costs 1.6× less per token, which makes it the better buy when GPT-5 Nano's lead doesn't matter for your workload.

Which is cheaper, GPT-5 Nano or Phi-4?

Phi-4 is cheaper. It lists at $0.07 per million input tokens and $0.14 per million output tokens; GPT-5 Nano lists at $0.05 and $0.40.

Is GPT-5 Nano or Phi-4 better for coding?

They score almost the same on coding (33.6 vs 34.4); test both on your own repository before choosing.

Which has the bigger context window?

GPT-5 Nano does, with 400K tokens against 128K.

How many benchmarks do GPT-5 Nano and Phi-4 share?

23 benchmarks have published results for both models. GPT-5 Nano has 49 scored results on Noometry and Phi-4 has 37.

Related comparisons

Go deeper