Model comparison

Gemma 4 31B IT vs Qwen3 Max

Gemma 4 31B IT and Qwen3 Max score almost the same on the Noometry Index (43.5 vs 43.7), so choose on price, context window or the category you care about most.

Last verified . 24 shared benchmarks.

Gemma 4 31B IT Google

43.5

Rank #90 Confirmed

Qwen3 Max Alibaba (Qwen)

43.7

Rank #87 Confirmed

Summary

  • They share 24 benchmarks with published results for both. Gemma 4 31B IT scores higher in 5 categories and Qwen3 Max in 3 categories; 5 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Qwen3 Max leads 48.1 to 37.9.
  • The biggest single-benchmark swing is NYT Connections (extended): 70.6% for Gemma 4 31B IT and 30.1% for Qwen3 Max.
  • Gemma 4 31B IT is cheaper at $0.09 / $0.34 per million input/output tokens, against $1.20 / $6 for Qwen3 Max.
  • Gemma 4 31B IT has downloadable open weights; the other is API-only.

Side by side

Gemma 4 31B IT and Qwen3 Max specifications
Gemma 4 31B ITQwen3 Max
ProviderGoogleAlibaba (Qwen)
Noometry Index43.543.7
Released2026-04-022025-09-23
WeightsOpenProprietary
Context window262K262K
Max output33K66K
Input $ / M tokens$0.09$1.20
Output $ / M tokens$0.34$6
Results tracked3533

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

Gemma 4 31B IT: 42.3 (#108), Qwen3 Max: 43.0 (#93)

Coding benchmarks
BenchmarkGemma 4 31B ITQwen3 Max
LMArena Coding14591456
ALE-Bench925.5370.45
LMArena WebDev1366—
SciCode43.4%—
WeirdML52.3%—

Agentic & Tool Use Not comparable

Gemma 4 31B IT: —, Qwen3 Max: —

Agentic & Tool Use benchmarks
BenchmarkGemma 4 31B ITQwen3 Max
Vending-Bench 2—71.56

Reasoning Gemma 4 31B IT leads

Gemma 4 31B IT: 27.2 (#122), Qwen3 Max: 22.6 (#190)

Reasoning benchmarks
BenchmarkGemma 4 31B ITQwen3 Max
Kagi LLM Benchmark63.5%72.5%
NYT Connections (extended)70.6%30.1%
Chess Puzzles5%4%
LMArena Hard Prompts14481448
DTBench82.7%82.1%
LMCA39.3%28.3%
Epoch Capabilities Index142.74142.38
CritPt1.4%—
Thematic Generalization53%—
Mystery Game Puzzles—5%
Surface Evolver Bench30.6%—

Math Gemma 4 31B IT leads

Gemma 4 31B IT: 43.2 (#81), Qwen3 Max: 38.7 (#131)

Math benchmarks
BenchmarkGemma 4 31B ITQwen3 Max
OTIS Mock AIME 2024-202573.3%73.3%
LMArena Math14651446
FrontierMath (Tiers 1-3)—18.9%
MATH Level 5—97.1%

Knowledge Qwen3 Max leads

Gemma 4 31B IT: 37.9 (#151), Qwen3 Max: 48.1 (#78)

Knowledge benchmarks
BenchmarkGemma 4 31B ITQwen3 Max
GPQA Diamond75.8%72.6%
SimpleQA Verified10.4%48.7%
LMArena Expert14651455
Vectara Hallucination Rate7.4%—

Multimodal Not comparable

Gemma 4 31B IT: 41.6 (#34), Qwen3 Max: —

Multimodal benchmarks
BenchmarkGemma 4 31B ITQwen3 Max
LMArena Vision1277—
LMArena Document1425—

Multilingual Too close to call

Gemma 4 31B IT: 53.8 (#57), Qwen3 Max: 53.7 (#62)

Multilingual benchmarks
BenchmarkGemma 4 31B ITQwen3 Max
LMArena Non-English14311429
LMArena Chinese14761478
LMArena French14351449
LMArena Russian14601428
LMArena Spanish14441462
LMArena German—1463
LMArena Japanese—1397
LMArena Korean—1399

Instruction Following Too close to call

Gemma 4 31B IT: 75.5 (#61), Qwen3 Max: 74.8 (#87)

Instruction Following benchmarks
BenchmarkGemma 4 31B ITQwen3 Max
LMArena Instruction Following14331419

Long Context Gemma 4 31B IT leads

Gemma 4 31B IT: 44.2 (#71), Qwen3 Max: 41.6 (#134)

Long Context benchmarks
BenchmarkGemma 4 31B ITQwen3 Max
LMArena Longer Query14461438
Fiction.LiveBench—66.7%
CL-bench—14.5%

Writing & Preference Qwen3 Max leads

Gemma 4 31B IT: 60.5 (#96), Qwen3 Max: 62.4 (#76)

Writing & Preference benchmarks
BenchmarkGemma 4 31B ITQwen3 Max
LMArena Text14431439
LMArena Creative Writing14151402
LMArena Multi-Turn14521446
EQ-Bench Creative Writing1368—
EQ-Bench 41120—

Frequently asked questions

Is Gemma 4 31B IT better than Qwen3 Max?

Gemma 4 31B IT and Qwen3 Max score almost the same on the Noometry Index (43.5 vs 43.7), so choose on price, context window or the category you care about most.

Which is cheaper, Gemma 4 31B IT or Qwen3 Max?

Gemma 4 31B IT is cheaper. It lists at $0.09 per million input tokens and $0.34 per million output tokens; Qwen3 Max lists at $1.20 and $6.

Is Gemma 4 31B IT or Qwen3 Max better for coding?

They score almost the same on coding (42.3 vs 43.0); test both on your own repository before choosing.

Which has the bigger context window?

Both accept 262K tokens.

How many benchmarks do Gemma 4 31B IT and Qwen3 Max share?

24 benchmarks have published results for both models. Gemma 4 31B IT has 35 scored results on Noometry and Qwen3 Max has 33.

Related comparisons

Go deeper