Model comparison

Gemma 4 31B IT vs Grok 4.3

Gemma 4 31B IT and Grok 4.3 score almost the same on the Noometry Index (43.5 vs 43.8), so choose on price, context window or the category you care about most.

Last verified . 29 shared benchmarks.

Gemma 4 31B IT Google

43.5

Rank #90 Confirmed

Grok 4.3 xAI

43.8

Rank #86 Confirmed

Summary

  • They share 29 benchmarks with published results for both. Gemma 4 31B IT scores higher in 6 categories and Grok 4.3 in 3 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Grok 4.3 leads 52.5 to 37.9.
  • The biggest single-benchmark swing is SimpleQA Verified: 10.4% for Gemma 4 31B IT and 33.2% for Grok 4.3.
  • Gemma 4 31B IT is cheaper at $0.09 / $0.34 per million input/output tokens, against $1.25 / $2.50 for Grok 4.3.
  • Grok 4.3 accepts more context: 1M tokens versus 262K.
  • Gemma 4 31B IT has downloadable open weights; the other is API-only.

Side by side

Gemma 4 31B IT and Grok 4.3 specifications
Gemma 4 31B ITGrok 4.3
ProviderGooglexAI
Noometry Index43.543.8
Released2026-04-022026-04-17
WeightsOpenProprietary
Context window262K1M
Max output33K30K
Input $ / M tokens$0.09$1.25
Output $ / M tokens$0.34$2.50
Results tracked3540

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

Gemma 4 31B IT: 42.3 (#108), Grok 4.3: 41.6 (#121)

Coding benchmarks
BenchmarkGemma 4 31B ITGrok 4.3
LMArena WebDev13661357
SciCode43.4%47.3%
WeirdML52.3%49.9%
LMArena Coding14591415
ALE-Bench925.5944.17

Agentic & Tool Use Not comparable

Gemma 4 31B IT: —, Grok 4.3: 27.7 (#99)

Agentic & Tool Use benchmarks
BenchmarkGemma 4 31B ITGrok 4.3
GDP.pdf—8%
LMArena Search—1165
Vending-Bench 2—35.26

Reasoning Grok 4.3 leads

Gemma 4 31B IT: 27.2 (#122), Grok 4.3: 35.9 (#68)

Reasoning benchmarks
BenchmarkGemma 4 31B ITGrok 4.3
NYT Connections (extended)70.6%55.2%
CritPt1.4%8%
Chess Puzzles5%25%
LMArena Hard Prompts14481396
DTBench82.7%90.7%
LMCA39.3%38.3%
Epoch Capabilities Index142.74149.16
Kagi LLM Benchmark63.5%—
Thematic Generalization53%—
Surface Evolver Bench30.6%—
ForecastBench—60.3

Math Grok 4.3 leads

Gemma 4 31B IT: 43.2 (#81), Grok 4.3: 46.0 (#74)

Math benchmarks
BenchmarkGemma 4 31B ITGrok 4.3
OTIS Mock AIME 2024-202573.3%93.3%
LMArena Math14651388
FrontierMath (Tiers 1-3)—42.8%
FrontierMath Tier 4—14.6%
ProofBench—11%

Knowledge Grok 4.3 leads

Gemma 4 31B IT: 37.9 (#151), Grok 4.3: 52.5 (#62)

Knowledge benchmarks
BenchmarkGemma 4 31B ITGrok 4.3
GPQA Diamond75.8%88.8%
SimpleQA Verified10.4%33.2%
LMArena Expert14651385
Vectara Hallucination Rate7.4%—

Multimodal Gemma 4 31B IT leads

Gemma 4 31B IT: 41.6 (#34), Grok 4.3: 31.6 (#104)

Multimodal benchmarks
BenchmarkGemma 4 31B ITGrok 4.3
LMArena Vision12771229
Blueprint-Bench 2—0%
LMArena Document1425—

Multilingual Gemma 4 31B IT leads

Gemma 4 31B IT: 53.8 (#57), Grok 4.3: 50.5 (#120)

Multilingual benchmarks
BenchmarkGemma 4 31B ITGrok 4.3
LMArena Non-English14311385
LMArena Chinese14761422
LMArena French14351412
LMArena Russian14601399
LMArena Spanish14441398
LMArena German—1395
LMArena Japanese—1379
LMArena Korean—1356

Instruction Following Gemma 4 31B IT leads

Gemma 4 31B IT: 75.5 (#61), Grok 4.3: 72.1 (#140)

Instruction Following benchmarks
BenchmarkGemma 4 31B ITGrok 4.3
LMArena Instruction Following14331366

Long Context Gemma 4 31B IT leads

Gemma 4 31B IT: 44.2 (#71), Grok 4.3: 42.5 (#123)

Long Context benchmarks
BenchmarkGemma 4 31B ITGrok 4.3
LMArena Longer Query14461393

Writing & Preference Gemma 4 31B IT leads

Gemma 4 31B IT: 60.5 (#96), Grok 4.3: 58.5 (#118)

Writing & Preference benchmarks
BenchmarkGemma 4 31B ITGrok 4.3
LMArena Text14431397
LMArena Creative Writing14151380
EQ-Bench 411201075
LMArena Multi-Turn14521406
EQ-Bench Creative Writing1368—

Frequently asked questions

Is Gemma 4 31B IT better than Grok 4.3?

Gemma 4 31B IT and Grok 4.3 score almost the same on the Noometry Index (43.5 vs 43.8), so choose on price, context window or the category you care about most.

Which is cheaper, Gemma 4 31B IT or Grok 4.3?

Gemma 4 31B IT is cheaper. It lists at $0.09 per million input tokens and $0.34 per million output tokens; Grok 4.3 lists at $1.25 and $2.50.

Is Gemma 4 31B IT or Grok 4.3 better for coding?

They score almost the same on coding (42.3 vs 41.6); test both on your own repository before choosing.

Which has the bigger context window?

Grok 4.3 does, with 1M tokens against 262K.

How many benchmarks do Gemma 4 31B IT and Grok 4.3 share?

29 benchmarks have published results for both models. Gemma 4 31B IT has 35 scored results on Noometry and Grok 4.3 has 40.

Related comparisons

Go deeper