Microsoft, open weights

Phi-4

Phi-4 by Microsoft ranks 279th of 354 ranked models on the Noometry Index as of October 2026, with a score of 31.2. Its strongest category is agentic & tool use, where it ranks 128th. API pricing starts at $0.07 per million input tokens and $0.14 per million output tokens, with a 128K-token context window.

Last verified

Specifications

Noometry rank
#279 of 354
Index score
31.2
Evidence
Confirmed 37 results
Provider
Microsoft
Released
December 11, 2024
Weights
Open weights
Reasoning
No
Context window
128K
Max output
4K
Input price
$0.07 / M
Output price
$0.14 / M
Blended price
$0.0875 / M
Output speed
Not measured
Value
#15 of 219
Knowledge cutoff
October 2023
Input
text
Hugging Face
microsoft/phi-4

Category scores

Each category score combines every public result we have in that category.

Phi-4 category scores
  1. Coding 34.4
  2. Agentic & Tool Use 22.8
  3. Reasoning 17.7
  4. Math 20.8
  5. Knowledge 32.6
  6. Multilingual 37.2
  7. Instruction Following 60.4
  8. Long Context 36.9
  9. Writing & Preference 40.5
Phi-4 category ranks
CategoryScoreRankResults
Coding34.4#2394
Agentic & Tool Use22.8#1282
Reasoning17.7#2914
Math20.8#2854
Knowledge32.6#2094
Multilingual37.2#2371
Instruction Following60.4#2512
Long Context36.9#2261
Writing & Preference40.5#2445

Strengths and weaknesses

Categories where Phi-4 places highest and lowest among the models ranked in each, with its score against that category's median.

Strongest categories

Phi-4: strongest categories
CategoryScorevs medianRank
Knowledge32.6−4.7#209 of 314, top 67%
Coding34.4−4.4#239 of 340, top 71%
Long Context36.9−4.0#226 of 296, top 77%

Weakest categories

Phi-4: weakest categories
CategoryScorevs medianRank
Math20.8−15.8#285 of 327, top 88%
Reasoning17.7−5.9#291 of 350, top 84%
Agentic & Tool Use22.8−7.6#128 of 154, top 84%

Closest competitors

The models ranked just above and below Phi-4. When scores are this close, price and speed are often the better way to choose.

Models ranked closest to Phi-4
ModelRankScoreBlended $/MSpeed
Qwen-14B#27531.4——Compare
Phi 3 Mini 4k Instruct June 2024#27631.3——Compare
Gemma 1.1 7b IT#27731.3——Compare
Mistral Small 3#27831.2$0.0575—Compare
Mistral Small 3.2#28031.2$0.1368Compare
Amazon Nova Pro#28131.0$1.40—Compare
Llama 4 Maverick#28230.9$0.30456Compare
Phi-4 Mini#28330.9$0.13—Compare

Sponsored placements are available on pages like this one. Advertise on Noometry

Benchmark results

Every published result we track, with its source. Bold rows are the ones used for ranking; where several exist we prefer independent runs over self-reported numbers.

Coding

Phi-4 Coding benchmark results
BenchmarkScorePositionSettingSourceDate
BigCodeBench Instruct45.5%#19 of 64, top 30%BigCodeBench2024-12-13
LiveBench Coding30.7%#31 of 39, top 80%Epoch AI
LMArena Coding1231#230 of 294, top 79%LMArena2026-10-08
BigCodeBench Complete55.4%#17 of 66, top 26%BigCodeBench2024-12-13

Agentic & Tool Use

Phi-4 Agentic & Tool Use benchmark results
BenchmarkScorePositionSettingSourceDate
Berkeley Function Calling Leaderboard28.8%#35 of 49, top 72%promptBerkeley Function Calling Leaderboard
BALROG11.6%#32 of 35, top 92%Epoch AI

Reasoning

Phi-4 Reasoning benchmark results
BenchmarkScorePositionSettingSourceDate
Chess Puzzles1%#107 of 129, top 83%Epoch AI2026-08-28
LiveBench Reasoning47.8%#21 of 39, top 54%Epoch AI
LMArena Hard Prompts1220#231 of 297, top 78%LMArena2026-10-08
LiveBench Data Analysis45.2%#30 of 39, top 77%Epoch AI
Epoch Capabilities Index130.42#137 of 213, top 65%Epoch AI2024-12-12
LiveBench41.6%#29 of 39, top 75%Epoch AI

Math

Phi-4 Math benchmark results
BenchmarkScorePositionSettingSourceDate
OTIS Mock AIME 2024-202513.8%#131 of 173, top 76%Epoch AI2025-02-25
LiveBench Math42%#25 of 39, top 65%Epoch AI
LMArena Math1246#216 of 285, top 76%LMArena2026-10-08
MATH Level 564.9%#35 of 79, top 45%Epoch AI2025-01-31

Knowledge

Phi-4 Knowledge benchmark results
BenchmarkScorePositionSettingSourceDate
GPQA Diamond56.1%#120 of 186, top 65%Epoch AI2025-01-31
Confabulations (lower is better)29.4%#46 of 51, top 91%Lech Mazur benchmarks
Vectara Hallucination Rate (lower is better)3.7%#3 of 96, top 4%Vectara Hallucination Leaderboard
LMArena Expert1203#220 of 273, top 81%LMArena2026-10-08
MMLU84.8%#8 of 81, top 10%Epoch AI

Multilingual

Phi-4 Multilingual benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Non-English1197#237 of 297, top 80%LMArena2026-10-08
LMArena Chinese1212#228 of 285, top 80%LMArena2026-10-08
LMArena French1224#190 of 223, top 86%LMArena2026-10-08
LMArena German1222#185 of 231, top 81%LMArena2026-10-08
LMArena Japanese1158#172 of 211, top 82%LMArena2026-10-08
LMArena Korean1151#177 of 213, top 84%LMArena2026-10-08
LMArena Russian1209#231 of 283, top 82%LMArena2026-10-08
LMArena Spanish1234#188 of 226, top 84%LMArena2026-10-08

Instruction Following

Phi-4 Instruction Following benchmark results
BenchmarkScorePositionSettingSourceDate
LiveBench Instruction Following58.4%#29 of 39, top 75%Epoch AI
LMArena Instruction Following1201#234 of 298, top 79%LMArena2026-10-08

Long Context

Phi-4 Long Context benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Longer Query1217#236 of 291, top 82%LMArena2026-10-08

Writing & Preference

Phi-4 Writing & Preference benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Text1217#240 of 297, top 81%LMArena2026-10-08
LMArena Creative Writing1182#239 of 295, top 82%LMArena2026-10-08
Short-Story Creative Writing62.6%#36 of 39, top 93%Epoch AI
LMArena Multi-Turn1206#236 of 295, top 80%LMArena2026-10-08
LiveBench Language25.6%#31 of 39, top 80%Epoch AI

API pricing by provider

Phi-4 API prices
RouteInput $/MOutput $/MCached input $/MChecked
azure$0.13$0.50—2026-10-10
openrouter$0.07$0.14—2026-10-10

Compare Phi-4

Other Microsoft models

Frequently asked questions

How good is Phi-4?

Phi-4 by Microsoft ranks 279th of 354 ranked models on the Noometry Index as of October 2026, with a score of 31.2. Its strongest category is agentic & tool use, where it ranks 128th. API pricing starts at $0.07 per million input tokens and $0.14 per million output tokens, with a 128K-token context window.

How much does Phi-4 cost?

Phi-4 costs $0.07 per million input tokens and $0.14 per million output tokens on openrouter.

What is Phi-4's context window?

Phi-4 accepts up to 128K tokens of input and can write up to 4K tokens in one response.

Is Phi-4 open source?

Yes. Phi-4's weights are downloadable from Hugging Face (microsoft/phi-4); check the license for commercial terms.

What are Phi-4's strengths and weaknesses?

Relative to other ranked models, Phi-4 places best in knowledge, coding, long context and lowest in math, reasoning, agentic & tool use.

What is Phi-4 best at?

Its best category is agentic & tool use, where it ranks 128th on Noometry.