Model comparison

Granite 4.0 Micro vs Phi 3 Small 8k Instruct

Granite 4.0 Micro and Phi 3 Small 8k Instruct score almost the same on the Noometry Index (29.0 vs 29.3), so choose on price, context window or the category you care about most.

Last verified . 0 shared benchmarks.

Granite 4.0 Micro IBM

29.0

Rank #318 Confirmed

Phi 3 Small 8k Instruct Microsoft

29.3

Rank #314 Confirmed

Summary

  • The widest gap is in knowledge, where Phi 3 Small 8k Instruct leads 29.1 to 9.9.

Side by side

Granite 4.0 Micro and Phi 3 Small 8k Instruct specifications
Granite 4.0 MicroPhi 3 Small 8k Instruct
ProviderIBMMicrosoft
Noometry Index29.029.3
Released2025-10-022024-04-23
WeightsOpenOpen
Context window131K—
Max output118K—
Input $ / M tokens$0.017—
Output $ / M tokens$0.11—
Results tracked832

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Not comparable

Granite 4.0 Micro: —, Phi 3 Small 8k Instruct: 27.9 (#318)

Coding benchmarks
BenchmarkGranite 4.0 MicroPhi 3 Small 8k Instruct
LiveBench Coding—20.3%
LMArena Coding—1101

Reasoning Granite 4.0 Micro leads

Granite 4.0 Micro: 19.2 (#265), Phi 3 Small 8k Instruct: 14.8 (#323)

Reasoning benchmarks
BenchmarkGranite 4.0 MicroPhi 3 Small 8k Instruct
Chess Puzzles0%—
LiveBench Reasoning—15.9%
LMArena Hard Prompts—1100
LiveBench Data Analysis—30.3%
Adversarial NLI—58.1%
BIG-Bench Hard—79.1%
HellaSwag—77%
LiveBench—24%
WinoGrande—81.5%

Math Phi 3 Small 8k Instruct leads

Granite 4.0 Micro: 12.0 (#307), Phi 3 Small 8k Instruct: 27.6 (#248)

Math benchmarks
BenchmarkGranite 4.0 MicroPhi 3 Small 8k Instruct
OTIS Mock AIME 2024-20252.8%—
Omni-MATH20.9%—
LiveBench Math—17.6%
LMArena Math—1151

Knowledge Phi 3 Small 8k Instruct leads

Granite 4.0 Micro: 9.9 (#304), Phi 3 Small 8k Instruct: 29.1 (#240)

Knowledge benchmarks
BenchmarkGranite 4.0 MicroPhi 3 Small 8k Instruct
GPQA Diamond28.3%—
MMLU-Pro39.5%—
GPQA (HELM)30.7%—
LMArena Expert—1067
ARC (AI2) Challenge—90.7%
MMLU—75.7%
OpenBookQA—88%
TriviaQA—58.1%

Multilingual Not comparable

Granite 4.0 Micro: —, Phi 3 Small 8k Instruct: 28.5 (#272)

Multilingual benchmarks
BenchmarkGranite 4.0 MicroPhi 3 Small 8k Instruct
LMArena Non-English—1058
LMArena Chinese—1061
LMArena French—1135
LMArena German—1080
LMArena Japanese—966
LMArena Korean—894
LMArena Russian—1111
LMArena Spanish—1111

Instruction Following Granite 4.0 Micro leads

Granite 4.0 Micro: 69.9 (#169), Phi 3 Small 8k Instruct: 51.9 (#292)

Instruction Following benchmarks
BenchmarkGranite 4.0 MicroPhi 3 Small 8k Instruct
LiveBench Instruction Following—47.2%
IFEval84.9%—
LMArena Instruction Following—1087

Long Context Not comparable

Granite 4.0 Micro: —, Phi 3 Small 8k Instruct: 33.0 (#267)

Long Context benchmarks
BenchmarkGranite 4.0 MicroPhi 3 Small 8k Instruct
LMArena Longer Query—1088

Writing & Preference Granite 4.0 Micro leads

Granite 4.0 Micro: 46.7 (#216), Phi 3 Small 8k Instruct: 31.1 (#284)

Writing & Preference benchmarks
BenchmarkGranite 4.0 MicroPhi 3 Small 8k Instruct
LMArena Text—1110
LMArena Creative Writing—1083
WildBench67%—
LMArena Multi-Turn—1068
LiveBench Language—12.9%

Frequently asked questions

Is Granite 4.0 Micro better than Phi 3 Small 8k Instruct?

Granite 4.0 Micro and Phi 3 Small 8k Instruct score almost the same on the Noometry Index (29.0 vs 29.3), so choose on price, context window or the category you care about most.

How many benchmarks do Granite 4.0 Micro and Phi 3 Small 8k Instruct share?

0 benchmarks have published results for both models. Granite 4.0 Micro has 8 scored results on Noometry and Phi 3 Small 8k Instruct has 32.

Related comparisons

Go deeper