Model comparison

Granite 3.0 8b Instruct vs Phi-4 Mini

Granite 3.0 8b Instruct and Phi-4 Mini score almost the same on the Noometry Index (31.6 vs 30.9), so choose on price, context window or the category you care about most.

Last verified . 0 shared benchmarks.

Granite 3.0 8b Instruct IBM

31.6

Rank #270 Confirmed

Phi-4 Mini Microsoft

30.9

Rank #283 Reported

Summary

  • The widest gap is in knowledge, where Granite 3.0 8b Instruct leads 29.6 to 25.3.

Side by side

Granite 3.0 8b Instruct and Phi-4 Mini specifications
Granite 3.0 8b InstructPhi-4 Mini
ProviderIBMMicrosoft
Noometry Index31.630.9
Released—2024-12-11
WeightsOpenOpen
Context window—128K
Max output—4K
Input $ / M tokens—$0.075
Output $ / M tokens—$0.30
Results tracked143

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Granite 3.0 8b Instruct leads

Granite 3.0 8b Instruct: 29.7 (#301), Phi-4 Mini: 28.1 (#317)

Coding benchmarks
BenchmarkGranite 3.0 8b InstructPhi-4 Mini
SciCode—10.8%
BigCodeBench Instruct29.3%—
LMArena Coding1112—
BigCodeBench Complete35.4%—

Reasoning Phi-4 Mini leads

Granite 3.0 8b Instruct: 20.9 (#230), Phi-4 Mini: 22.4 (#195)

Reasoning benchmarks
BenchmarkGranite 3.0 8b InstructPhi-4 Mini
CritPt—0%
LMArena Hard Prompts1092—

Math Not comparable

Granite 3.0 8b Instruct: 32.8 (#210), Phi-4 Mini: —

Math benchmarks
BenchmarkGranite 3.0 8b InstructPhi-4 Mini
LMArena Math1143—

Knowledge Granite 3.0 8b Instruct leads

Granite 3.0 8b Instruct: 29.6 (#236), Phi-4 Mini: 25.3 (#262)

Knowledge benchmarks
BenchmarkGranite 3.0 8b InstructPhi-4 Mini
Vectara Hallucination Rate—23.5%
LMArena Expert1087—

Multilingual Not comparable

Granite 3.0 8b Instruct: 27.2 (#276), Phi-4 Mini: —

Multilingual benchmarks
BenchmarkGranite 3.0 8b InstructPhi-4 Mini
LMArena Non-English1037—
LMArena Chinese1063—
LMArena Russian1060—

Instruction Following Not comparable

Granite 3.0 8b Instruct: 56.0 (#276), Phi-4 Mini: —

Instruction Following benchmarks
BenchmarkGranite 3.0 8b InstructPhi-4 Mini
LMArena Instruction Following1088—

Long Context Not comparable

Granite 3.0 8b Instruct: 34.0 (#252), Phi-4 Mini: —

Long Context benchmarks
BenchmarkGranite 3.0 8b InstructPhi-4 Mini
LMArena Longer Query1122—

Writing & Preference Not comparable

Granite 3.0 8b Instruct: 31.1 (#285), Phi-4 Mini: —

Writing & Preference benchmarks
BenchmarkGranite 3.0 8b InstructPhi-4 Mini
LMArena Text1096—
LMArena Creative Writing1071—
LMArena Multi-Turn1063—

Frequently asked questions

Is Granite 3.0 8b Instruct better than Phi-4 Mini?

Granite 3.0 8b Instruct and Phi-4 Mini score almost the same on the Noometry Index (31.6 vs 30.9), so choose on price, context window or the category you care about most.

Is Granite 3.0 8b Instruct or Phi-4 Mini better for coding?

Granite 3.0 8b Instruct scores higher on coding benchmarks: 29.7 versus 28.1 in the Noometry coding category.

How many benchmarks do Granite 3.0 8b Instruct and Phi-4 Mini share?

0 benchmarks have published results for both models. Granite 3.0 8b Instruct has 14 scored results on Noometry and Phi-4 Mini has 3.

Related comparisons

Go deeper