Microsoft, open weights
Phi 3 Mini 4k Instruct
Phi 3 Mini 4k Instruct by Microsoft ranks 328th of 354 ranked models on the Noometry Index as of October 2026, with a score of 27.9. Its strongest category is knowledge, where it ranks 246th.
Last verified
Specifications
- Noometry rank
- #328 of 354
- Index score
- 27.9
- Evidence
- Confirmed 35 results
- Provider
Microsoft
- Released
- April 23, 2024
- Weights
- Open weights
- Reasoning
- Unknown
- Context window
- —
- Max output
- —
- Input price
- Not listed
- Output price
- Not listed
- Blended price
- Not listed
- Output speed
- Not measured
- Value
- Not ranked
- Knowledge cutoff
- Unknown
Category scores
Each category score combines every public result we have in that category.
- Coding 26.6
- Reasoning 14.1
- Math 26.6
- Knowledge 28.5
- Multilingual 26.3
- Instruction Following 47.7
- Long Context 31.7
- Writing & Preference 27.6
| Category | Score | Rank | Results |
|---|---|---|---|
| Coding | 26.6 | #323 | 2 |
| Reasoning | 14.1 | #328 | 4 |
| Math | 26.6 | #257 | 2 |
| Knowledge | 28.5 | #246 | 1 |
| Multilingual | 26.3 | #280 | 1 |
| Instruction Following | 47.7 | #303 | 2 |
| Long Context | 31.7 | #276 | 1 |
| Writing & Preference | 27.6 | #300 | 4 |
Strengths and weaknesses
Categories where Phi 3 Mini 4k Instruct places highest and lowest among the models ranked in each, with its score against that category's median.
Strongest categories
| Category | Score | vs median | Rank |
|---|---|---|---|
| Knowledge | 28.5 | −8.8 | #246 of 314, top 79% |
| Math | 26.6 | −9.9 | #257 of 327, top 79% |
| Long Context | 31.7 | −9.2 | #276 of 296, top 94% |
Weakest categories
| Category | Score | vs median | Rank |
|---|---|---|---|
| Instruction Following | 47.7 | −23.5 | #303 of 305, top 100% |
| Writing & Preference | 27.6 | −26.2 | #300 of 312, top 97% |
| Coding | 26.6 | −12.1 | #323 of 340, top 95% |
Closest competitors
The models ranked just above and below Phi 3 Mini 4k Instruct. When scores are this close, price and speed are often the better way to choose.
| Model | Rank | Score | Blended $/M | Speed | |
|---|---|---|---|---|---|
| GPT-4o | #324 | 28.6 | $4.38 | — | Compare |
| Ministral 8B | #325 | 28.2 | $0.15 | — | Compare |
| Gemma 3 4B | #326 | 28.1 | $0.05 | 72 | Compare |
| GPT-4.1 nano | #327 | 27.9 | $0.18 | 135 | Compare |
| Yi-34B | #329 | 27.8 | — | — | Compare |
| Llama 4 Scout | #330 | 27.7 | $0.15 | 272 | Compare |
| Llama 3.2 90B | #331 | 27.5 | — | — | Compare |
| Gemini 1.0 Pro | #332 | 27.3 | — | — | Compare |
Sponsored placements are available on pages like this one. Advertise on Noometry
Benchmark results
Every published result we track, with its source. Bold rows are the ones used for ranking; where several exist we prefer independent runs over self-reported numbers.
Coding
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LiveBench Coding | 15.5% | #38 of 39, top 98% | Epoch AI | ||
| LMArena Coding | 1093 | #273 of 294, top 93% | LMArena | 2026-10-08 | |
| HumanEval+ | 59.1% | #29 of 45, top 65% | EvalPlus | ||
| MBPP+ | 54.2% | #32 of 38, top 85% | EvalPlus |
Reasoning
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| Chess Puzzles | 0% | #125 of 129, top 97% | Epoch AI | 2026-08-30 | |
| LiveBench Reasoning | 26.8% | #33 of 39, top 85% | Epoch AI | ||
| LMArena Hard Prompts | 1072 | #275 of 297, top 93% | LMArena | 2026-10-08 | |
| LiveBench Data Analysis | 34.7% | #35 of 39, top 90% | Epoch AI | ||
| Adversarial NLI | 52.8% | #6 of 9, top 67% | Epoch AI | ||
| BIG-Bench Hard | 71.7% | #9 of 27, top 34% | Epoch AI | ||
| HellaSwag | 76.7% | #23 of 29, top 80% | Epoch AI | ||
| LiveBench | 22.4% | #38 of 39, top 98% | Epoch AI | ||
| WinoGrande | 70.8% | #31 of 43, top 73% | Epoch AI |
Math
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LiveBench Math | 15.7% | #38 of 39, top 98% | Epoch AI | ||
| LMArena Math | 1111 | #264 of 285, top 93% | LMArena | 2026-10-08 |
Knowledge
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Expert | 1045 | #263 of 273, top 97% | LMArena | 2026-10-08 | |
| ARC (AI2) Challenge | 84.9% | #10 of 39, top 26% | Epoch AI | ||
| MMLU | 68.8% | #49 of 81, top 61% | Epoch AI | ||
| OpenBookQA | 88% | Best of 19 | Epoch AI | ||
| TriviaQA | 64% | #22 of 25, top 88% | Epoch AI |
Multilingual
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Non-English | 1021 | #280 of 297, top 95% | LMArena | 2026-10-08 | |
| LMArena Chinese | 1021 | #272 of 285, top 96% | LMArena | 2026-10-08 | |
| LMArena French | 1076 | #217 of 223, top 98% | LMArena | 2026-10-08 | |
| LMArena German | 1044 | #220 of 231, top 96% | LMArena | 2026-10-08 | |
| LMArena Japanese | 935 | #206 of 211, top 98% | LMArena | 2026-10-08 | |
| LMArena Korean | 905 | #209 of 213, top 99% | LMArena | 2026-10-08 | |
| LMArena Russian | 1022 | #271 of 283, top 96% | LMArena | 2026-10-08 | |
| LMArena Spanish | 1085 | #220 of 226, top 98% | LMArena | 2026-10-08 |
Instruction Following
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LiveBench Instruction Following | 39.1% | #39 of 39, top 100% | Epoch AI | ||
| LMArena Instruction Following | 1053 | #281 of 298, top 95% | LMArena | 2026-10-08 |
Long Context
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Longer Query | 1044 | #280 of 291, top 97% | LMArena | 2026-10-08 |
Writing & Preference
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Text | 1073 | #282 of 297, top 95% | LMArena | 2026-10-08 | |
| LMArena Creative Writing | 1037 | #281 of 295, top 96% | LMArena | 2026-10-08 | |
| LMArena Multi-Turn | 1018 | #284 of 295, top 97% | LMArena | 2026-10-08 | |
| LiveBench Language | 9.2% | #39 of 39, top 100% | Epoch AI |
Compare Phi 3 Mini 4k Instruct
- Phi 3 Mini 4k Instruct vs GPT-4.1 nano
- Phi 3 Mini 4k Instruct vs Yi-34B
- Phi 3 Mini 4k Instruct vs Gemma 3 4B
- Phi 3 Mini 4k Instruct vs Llama 4 Scout
- Phi 3 Mini 4k Instruct vs Ministral 8B
- Phi 3 Mini 4k Instruct vs Llama 3.2 90B
- Phi 3 Mini 4k Instruct vs GPT-6 Astra
- Phi 3 Mini 4k Instruct vs Claude Fable 5.1
- Phi 3 Mini 4k Instruct vs Gemini 3.8 Flash
- Phi 3 Mini 4k Instruct vs Kimi K3
- Phi 3 Mini 4k Instruct vs Grok 4.6
- Phi 3 Mini 4k Instruct vs Qwen3.8 Max
- Phi 3 Mini 4k Instruct vs GLM-5.3
- Phi 3 Mini 4k Instruct vs Muse Spark 1.3
Other Microsoft models
Frequently asked questions
How good is Phi 3 Mini 4k Instruct?
Phi 3 Mini 4k Instruct by Microsoft ranks 328th of 354 ranked models on the Noometry Index as of October 2026, with a score of 27.9. Its strongest category is knowledge, where it ranks 246th.
Is Phi 3 Mini 4k Instruct open source?
Yes. Phi 3 Mini 4k Instruct's weights are downloadable; check the license for commercial terms.
What are Phi 3 Mini 4k Instruct's strengths and weaknesses?
Relative to other ranked models, Phi 3 Mini 4k Instruct places best in knowledge, math, long context and lowest in instruction following, writing & preference, coding.
What is Phi 3 Mini 4k Instruct best at?
Its best category is knowledge, where it ranks 246th on Noometry.