Model comparison

Llama 3.2 90B vs Phi 3 Mini 4k Instruct

Llama 3.2 90B and Phi 3 Mini 4k Instruct score almost the same on the Noometry Index (27.5 vs 27.9), so choose on price, context window or the category you care about most.

Last verified . 1 shared benchmarks.

Llama 3.2 90B Meta

27.5

Rank #331 Confirmed

Phi 3 Mini 4k Instruct Microsoft

27.9

Rank #328 Confirmed

Summary

  • They share 1 benchmark with published results for both. Llama 3.2 90B scores higher in 1 category and Phi 3 Mini 4k Instruct in 2 categories; 3 gaps are clear of the uncertainty.
  • The widest gap is in math, where Phi 3 Mini 4k Instruct leads 26.6 to 11.1.

Side by side

Llama 3.2 90B and Phi 3 Mini 4k Instruct specifications
Llama 3.2 90BPhi 3 Mini 4k Instruct
ProviderMetaMicrosoft
Noometry Index27.527.9
Released2024-09-242024-04-23
WeightsOpenOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked935

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Not comparable

Llama 3.2 90B: —, Phi 3 Mini 4k Instruct: 26.6 (#323)

Coding benchmarks
BenchmarkLlama 3.2 90BPhi 3 Mini 4k Instruct
LiveBench Coding—15.5%
LMArena Coding—1093
HumanEval+—59.1%
MBPP+—54.2%

Agentic & Tool Use Not comparable

Llama 3.2 90B: 30.0 (#80), Phi 3 Mini 4k Instruct: —

Agentic & Tool Use benchmarks
BenchmarkLlama 3.2 90BPhi 3 Mini 4k Instruct
BALROG27.3%—

Reasoning Llama 3.2 90B leads

Llama 3.2 90B: 21.7 (#217), Phi 3 Mini 4k Instruct: 14.1 (#328)

Reasoning benchmarks
BenchmarkLlama 3.2 90BPhi 3 Mini 4k Instruct
Chess Puzzles—0%
EnigmaEval0.4%—
LiveBench Reasoning—26.8%
LMArena Hard Prompts—1072
LiveBench Data Analysis—34.7%
Adversarial NLI—52.8%
BIG-Bench Hard—71.7%
Epoch Capabilities Index125.5—
HellaSwag—76.7%
LiveBench—22.4%
WinoGrande—70.8%

Math Phi 3 Mini 4k Instruct leads

Llama 3.2 90B: 11.1 (#308), Phi 3 Mini 4k Instruct: 26.6 (#257)

Math benchmarks
BenchmarkLlama 3.2 90BPhi 3 Mini 4k Instruct
OTIS Mock AIME 2024-20252.6%—
LiveBench Math—15.7%
LMArena Math—1111
MATH Level 539.4%—

Knowledge Phi 3 Mini 4k Instruct leads

Llama 3.2 90B: 21.7 (#274), Phi 3 Mini 4k Instruct: 28.5 (#246)

Knowledge benchmarks
BenchmarkLlama 3.2 90BPhi 3 Mini 4k Instruct
MMLU80.3%68.8%
GPQA Diamond41%—
LMArena Expert—1045
ARC (AI2) Challenge—84.9%
OpenBookQA—88%
TriviaQA—64%

Multimodal Not comparable

Llama 3.2 90B: 25.4 (#124), Phi 3 Mini 4k Instruct: —

Multimodal benchmarks
BenchmarkLlama 3.2 90BPhi 3 Mini 4k Instruct
LMArena Vision1000—
GeoBench52%—

Multilingual Not comparable

Llama 3.2 90B: —, Phi 3 Mini 4k Instruct: 26.3 (#280)

Multilingual benchmarks
BenchmarkLlama 3.2 90BPhi 3 Mini 4k Instruct
LMArena Non-English—1021
LMArena Chinese—1021
LMArena French—1076
LMArena German—1044
LMArena Japanese—935
LMArena Korean—905
LMArena Russian—1022
LMArena Spanish—1085

Instruction Following Not comparable

Llama 3.2 90B: —, Phi 3 Mini 4k Instruct: 47.7 (#303)

Instruction Following benchmarks
BenchmarkLlama 3.2 90BPhi 3 Mini 4k Instruct
LiveBench Instruction Following—39.1%
LMArena Instruction Following—1053

Long Context Not comparable

Llama 3.2 90B: —, Phi 3 Mini 4k Instruct: 31.7 (#276)

Long Context benchmarks
BenchmarkLlama 3.2 90BPhi 3 Mini 4k Instruct
LMArena Longer Query—1044

Writing & Preference Not comparable

Llama 3.2 90B: —, Phi 3 Mini 4k Instruct: 27.6 (#300)

Writing & Preference benchmarks
BenchmarkLlama 3.2 90BPhi 3 Mini 4k Instruct
LMArena Text—1073
LMArena Creative Writing—1037
LMArena Multi-Turn—1018
LiveBench Language—9.2%

Frequently asked questions

Is Llama 3.2 90B better than Phi 3 Mini 4k Instruct?

Llama 3.2 90B and Phi 3 Mini 4k Instruct score almost the same on the Noometry Index (27.5 vs 27.9), so choose on price, context window or the category you care about most.

How many benchmarks do Llama 3.2 90B and Phi 3 Mini 4k Instruct share?

1 benchmark has published results for both models. Llama 3.2 90B has 9 scored results on Noometry and Phi 3 Mini 4k Instruct has 35.

Related comparisons

Go deeper