Model comparison
Llama 3.2 90B vs Phi 3 Mini 4k Instruct
Llama 3.2 90B and Phi 3 Mini 4k Instruct score almost the same on the Noometry Index (27.5 vs 27.9), so choose on price, context window or the category you care about most.
Last verified . 1 shared benchmarks.
Summary
- They share 1 benchmark with published results for both. Llama 3.2 90B scores higher in 1 category and Phi 3 Mini 4k Instruct in 2 categories; 3 gaps are clear of the uncertainty.
- The widest gap is in math, where Phi 3 Mini 4k Instruct leads 26.6 to 11.1.
Side by side
| Llama 3.2 90B | Phi 3 Mini 4k Instruct | |
|---|---|---|
| Provider | Meta | Microsoft |
| Noometry Index | 27.5 | 27.9 |
| Released | 2024-09-24 | 2024-04-23 |
| Weights | Open | Open |
| Context window | — | — |
| Max output | — | — |
| Input $ / M tokens | — | — |
| Output $ / M tokens | — | — |
| Results tracked | 9 | 35 |
Sponsored placements are available on pages like this one. Advertise on Noometry
Category by category
Coding Not comparable
Llama 3.2 90B: —, Phi 3 Mini 4k Instruct: 26.6 (#323)
| Benchmark | Llama 3.2 90B | Phi 3 Mini 4k Instruct |
|---|---|---|
| LiveBench Coding | — | 15.5% |
| LMArena Coding | — | 1093 |
| HumanEval+ | — | 59.1% |
| MBPP+ | — | 54.2% |
Agentic & Tool Use Not comparable
Llama 3.2 90B: 30.0 (#80), Phi 3 Mini 4k Instruct: —
| Benchmark | Llama 3.2 90B | Phi 3 Mini 4k Instruct |
|---|---|---|
| BALROG | 27.3% | — |
Reasoning Llama 3.2 90B leads
Llama 3.2 90B: 21.7 (#217), Phi 3 Mini 4k Instruct: 14.1 (#328)
| Benchmark | Llama 3.2 90B | Phi 3 Mini 4k Instruct |
|---|---|---|
| Chess Puzzles | — | 0% |
| EnigmaEval | 0.4% | — |
| LiveBench Reasoning | — | 26.8% |
| LMArena Hard Prompts | — | 1072 |
| LiveBench Data Analysis | — | 34.7% |
| Adversarial NLI | — | 52.8% |
| BIG-Bench Hard | — | 71.7% |
| Epoch Capabilities Index | 125.5 | — |
| HellaSwag | — | 76.7% |
| LiveBench | — | 22.4% |
| WinoGrande | — | 70.8% |
Math Phi 3 Mini 4k Instruct leads
Llama 3.2 90B: 11.1 (#308), Phi 3 Mini 4k Instruct: 26.6 (#257)
| Benchmark | Llama 3.2 90B | Phi 3 Mini 4k Instruct |
|---|---|---|
| OTIS Mock AIME 2024-2025 | 2.6% | — |
| LiveBench Math | — | 15.7% |
| LMArena Math | — | 1111 |
| MATH Level 5 | 39.4% | — |
Knowledge Phi 3 Mini 4k Instruct leads
Llama 3.2 90B: 21.7 (#274), Phi 3 Mini 4k Instruct: 28.5 (#246)
| Benchmark | Llama 3.2 90B | Phi 3 Mini 4k Instruct |
|---|---|---|
| MMLU | 80.3% | 68.8% |
| GPQA Diamond | 41% | — |
| LMArena Expert | — | 1045 |
| ARC (AI2) Challenge | — | 84.9% |
| OpenBookQA | — | 88% |
| TriviaQA | — | 64% |
Multimodal Not comparable
Llama 3.2 90B: 25.4 (#124), Phi 3 Mini 4k Instruct: —
| Benchmark | Llama 3.2 90B | Phi 3 Mini 4k Instruct |
|---|---|---|
| LMArena Vision | 1000 | — |
| GeoBench | 52% | — |
Multilingual Not comparable
Llama 3.2 90B: —, Phi 3 Mini 4k Instruct: 26.3 (#280)
| Benchmark | Llama 3.2 90B | Phi 3 Mini 4k Instruct |
|---|---|---|
| LMArena Non-English | — | 1021 |
| LMArena Chinese | — | 1021 |
| LMArena French | — | 1076 |
| LMArena German | — | 1044 |
| LMArena Japanese | — | 935 |
| LMArena Korean | — | 905 |
| LMArena Russian | — | 1022 |
| LMArena Spanish | — | 1085 |
Instruction Following Not comparable
Llama 3.2 90B: —, Phi 3 Mini 4k Instruct: 47.7 (#303)
| Benchmark | Llama 3.2 90B | Phi 3 Mini 4k Instruct |
|---|---|---|
| LiveBench Instruction Following | — | 39.1% |
| LMArena Instruction Following | — | 1053 |
Long Context Not comparable
Llama 3.2 90B: —, Phi 3 Mini 4k Instruct: 31.7 (#276)
| Benchmark | Llama 3.2 90B | Phi 3 Mini 4k Instruct |
|---|---|---|
| LMArena Longer Query | — | 1044 |
Writing & Preference Not comparable
Llama 3.2 90B: —, Phi 3 Mini 4k Instruct: 27.6 (#300)
| Benchmark | Llama 3.2 90B | Phi 3 Mini 4k Instruct |
|---|---|---|
| LMArena Text | — | 1073 |
| LMArena Creative Writing | — | 1037 |
| LMArena Multi-Turn | — | 1018 |
| LiveBench Language | — | 9.2% |
Frequently asked questions
Is Llama 3.2 90B better than Phi 3 Mini 4k Instruct?
Llama 3.2 90B and Phi 3 Mini 4k Instruct score almost the same on the Noometry Index (27.5 vs 27.9), so choose on price, context window or the category you care about most.
How many benchmarks do Llama 3.2 90B and Phi 3 Mini 4k Instruct share?
1 benchmark has published results for both models. Llama 3.2 90B has 9 scored results on Noometry and Phi 3 Mini 4k Instruct has 35.