Model comparison

Granite 4.2 30b vs Grok-3 mini

Granite 4.2 30b and Grok-3 mini score almost the same on the Noometry Index (41.8 vs 41.2), so choose on price, context window or the category you care about most.

Last verified . 11 shared benchmarks.

Granite 4.2 30b IBM

41.8

Rank #130 Confirmed

Grok-3 mini xAI

41.2

Rank #141 Confirmed

Summary

  • They share 11 benchmarks with published results for both. Granite 4.2 30b scores higher in 4 categories and Grok-3 mini in 3 categories; 4 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Granite 4.2 30b leads 27.8 to 13.6.
  • Granite 4.2 30b has downloadable open weights; the other is API-only.

Side by side

Granite 4.2 30b and Grok-3 mini specifications
Granite 4.2 30bGrok-3 mini
ProviderIBMxAI
Noometry Index41.841.2
Released—2025-04-09
WeightsOpenProprietary
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked1135

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

Granite 4.2 30b: 41.0 (#126), Grok-3 mini: 40.8 (#131)

Coding benchmarks
BenchmarkGranite 4.2 30bGrok-3 mini
LMArena Coding13961379
Aider Polyglot—49.3%
WeirdML—42.6%

Reasoning Granite 4.2 30b leads

Granite 4.2 30b: 27.8 (#112), Grok-3 mini: 13.6 (#334)

Reasoning benchmarks
BenchmarkGranite 4.2 30bGrok-3 mini
LMArena Hard Prompts13741375
ARC-AGI-2—0.4%
Kagi LLM Benchmark—61.3%
ARC-AGI-1—16.5%
Epoch Capabilities Index—140.35

Math Not comparable

Granite 4.2 30b: —, Grok-3 mini: 42.1 (#85)

Math benchmarks
BenchmarkGranite 4.2 30bGrok-3 mini
OTIS Mock AIME 2024-2025—77.8%
Omni-MATH—31.8%
LMArena Math—1386
MATH Level 5—90.9%
FrontierMath (Feb 2025 set)—5.9%

Knowledge Grok-3 mini leads

Granite 4.2 30b: 39.1 (#138), Grok-3 mini: 46.4 (#81)

Knowledge benchmarks
BenchmarkGranite 4.2 30bGrok-3 mini
LMArena Expert14061395
GPQA Diamond—76.3%
MMLU-Pro—79.9%
Confabulations—10.8%
GPQA (HELM)—67.5%

Multilingual Too close to call

Granite 4.2 30b: 47.3 (#151), Grok-3 mini: 48.1 (#145)

Multilingual benchmarks
BenchmarkGranite 4.2 30bGrok-3 mini
LMArena Non-English13401352
LMArena Chinese14141387
LMArena Russian13431353
LMArena French—1357
LMArena German—1349
LMArena Japanese—1342
LMArena Korean—1335
LMArena Spanish—1381

Instruction Following Grok-3 mini leads

Granite 4.2 30b: 71.2 (#155), Grok-3 mini: 78.5 (#9)

Instruction Following benchmarks
BenchmarkGranite 4.2 30bGrok-3 mini
LMArena Instruction Following13471357
IFEval—95.1%

Long Context Too close to call

Granite 4.2 30b: 41.4 (#140), Grok-3 mini: 41.0 (#147)

Long Context benchmarks
BenchmarkGranite 4.2 30bGrok-3 mini
LMArena Longer Query13591372
Fiction.LiveBench—66.7%

Writing & Preference Granite 4.2 30b leads

Granite 4.2 30b: 53.8 (#156), Grok-3 mini: 52.5 (#169)

Writing & Preference benchmarks
BenchmarkGranite 4.2 30bGrok-3 mini
LMArena Text13611370
LMArena Creative Writing12881342
LMArena Multi-Turn13391355
Short-Story Creative Writing—73.5%
WildBench—65.1%

Frequently asked questions

Is Granite 4.2 30b better than Grok-3 mini?

Granite 4.2 30b and Grok-3 mini score almost the same on the Noometry Index (41.8 vs 41.2), so choose on price, context window or the category you care about most.

Is Granite 4.2 30b or Grok-3 mini better for coding?

They score almost the same on coding (41.0 vs 40.8); test both on your own repository before choosing.

How many benchmarks do Granite 4.2 30b and Grok-3 mini share?

11 benchmarks have published results for both models. Granite 4.2 30b has 11 scored results on Noometry and Grok-3 mini has 35.

Related comparisons

Go deeper