Model comparison

Granite 4.2 30b vs o4-mini

Granite 4.2 30b and o4-mini score almost the same on the Noometry Index (41.8 vs 41.6), so choose on price, context window or the category you care about most.

Last verified . 11 shared benchmarks.

Granite 4.2 30b IBM

41.8

Rank #130 Confirmed

o4-mini OpenAI

41.6

Rank #132 Confirmed

Summary

  • They share 11 benchmarks with published results for both. Granite 4.2 30b scores higher in 3 categories and o4-mini in 4 categories; 4 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where o4-mini leads 43.6 to 39.1.
  • Granite 4.2 30b has downloadable open weights; the other is API-only.

Side by side

Granite 4.2 30b and o4-mini specifications
Granite 4.2 30bo4-mini
ProviderIBMOpenAI
Noometry Index41.841.6
Released—2025-04-16
WeightsOpenProprietary
Context window—200K
Max output—100K
Input $ / M tokens—$1.10
Output $ / M tokens—$4.40
Results tracked1160

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

Granite 4.2 30b: 41.0 (#126), o4-mini: 40.9 (#127)

Coding benchmarks
BenchmarkGranite 4.2 30bo4-mini
LMArena Coding13961368
SWE-bench Verified (bash only)—45%
Aider Polyglot—72%
GSO—3.6%
WeirdML—52.6%
CadEval—62%
ALE-Bench—826.17
AlgoTune—1.72

Agentic & Tool Use Not comparable

Granite 4.2 30b: —, o4-mini: 32.6 (#61)

Agentic & Tool Use benchmarks
BenchmarkGranite 4.2 30bo4-mini
Berkeley Function Calling Leaderboard—53.2%
GDPval—25.3%
METR Time Horizons—63.9%

Reasoning Granite 4.2 30b leads

Granite 4.2 30b: 27.8 (#112), o4-mini: 24.6 (#162)

Reasoning benchmarks
BenchmarkGranite 4.2 30bo4-mini
LMArena Hard Prompts13741351
ARC-AGI-2—6.1%
SimpleBench—38.7%
Kagi LLM Benchmark—67.6%
ARC-AGI-1—58.7%
CritPt—0.6%
Chess Puzzles—26%
EnigmaEval—9.2%
Mystery Game Puzzles—5%
DTBench—77.6%
LMCA—26.5%
Epoch Capabilities Index—145.64
ForecastBench—61.8

Math Not comparable

Granite 4.2 30b: —, o4-mini: 40.8 (#89)

Math benchmarks
BenchmarkGranite 4.2 30bo4-mini
FrontierMath (Tiers 1-3)—36.1%
FrontierMath Tier 4—4.9%
OTIS Mock AIME 2024-2025—81.7%
Omni-MATH—72%
LMArena Math—1389
MATH Level 5—97.8%
FrontierMath (Feb 2025 set)—24.8%
FrontierMath Tier 4 (v1)—6.3%

Knowledge o4-mini leads

Granite 4.2 30b: 39.1 (#138), o4-mini: 43.6 (#91)

Knowledge benchmarks
BenchmarkGranite 4.2 30bo4-mini
LMArena Expert14061343
GPQA Diamond—79.6%
Humanity's Last Exam—18.1%
SimpleQA Verified—19.6%
MMLU-Pro—82%
Confabulations—15.8%
Vectara Hallucination Rate—18.6%
GPQA (HELM)—73.5%

Multimodal Not comparable

Granite 4.2 30b: —, o4-mini: 40.2 (#49)

Multimodal benchmarks
BenchmarkGranite 4.2 30bo4-mini
LMArena Vision—1194
GeoBench—64%
VPCT—57.5%

Multilingual Too close to call

Granite 4.2 30b: 47.3 (#151), o4-mini: 47.0 (#154)

Multilingual benchmarks
BenchmarkGranite 4.2 30bo4-mini
LMArena Non-English13401337
LMArena Chinese14141354
LMArena Russian13431334
LMArena French—1364
LMArena German—1336
LMArena Japanese—1308
LMArena Korean—1312
LMArena Spanish—1347

Instruction Following o4-mini leads

Granite 4.2 30b: 71.2 (#155), o4-mini: 75.2 (#68)

Instruction Following benchmarks
BenchmarkGranite 4.2 30bo4-mini
LMArena Instruction Following13471321
IFEval—92.8%

Long Context o4-mini leads

Granite 4.2 30b: 41.4 (#140), o4-mini: 45.5 (#33)

Long Context benchmarks
BenchmarkGranite 4.2 30bo4-mini
LMArena Longer Query13591315
Fiction.LiveBench—77.8%

Writing & Preference Too close to call

Granite 4.2 30b: 53.8 (#156), o4-mini: 54.0 (#152)

Writing & Preference benchmarks
BenchmarkGranite 4.2 30bo4-mini
LMArena Text13611353
LMArena Creative Writing12881294
LMArena Multi-Turn13391350
Short-Story Creative Writing—75%
WildBench—85.4%

Frequently asked questions

Is Granite 4.2 30b better than o4-mini?

Granite 4.2 30b and o4-mini score almost the same on the Noometry Index (41.8 vs 41.6), so choose on price, context window or the category you care about most.

Is Granite 4.2 30b or o4-mini better for coding?

They score almost the same on coding (41.0 vs 40.9); test both on your own repository before choosing.

How many benchmarks do Granite 4.2 30b and o4-mini share?

11 benchmarks have published results for both models. Granite 4.2 30b has 11 scored results on Noometry and o4-mini has 60.

Related comparisons

Go deeper