Model comparison

Gemma 3 12B vs Gemma 3 27B

Gemma 3 12B is the stronger model overall, scoring 32.1 to 30.8 on the Noometry Index.

Last verified . 23 shared benchmarks.

Gemma 3 12B Google

32.1

Rank #262 Confirmed

Gemma 3 27B Google

30.8

Rank #284 Confirmed

Summary

  • They share 23 benchmarks with published results for both. Gemma 3 12B scores higher in 4 categories and Gemma 3 27B in 5 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in long context, where Gemma 3 12B leads 40.0 to 27.6.
  • The biggest single-benchmark swing is GPQA Diamond: 39.5% for Gemma 3 12B and 47.7% for Gemma 3 27B.
  • Gemma 3 12B is cheaper at $0.05 / $0.15 per million input/output tokens, against $0.08 / $0.16 for Gemma 3 27B.

Side by side

Gemma 3 12B and Gemma 3 27B specifications
Gemma 3 12BGemma 3 27B
ProviderGoogleGoogle
Noometry Index32.130.8
Released2025-03-122025-03-11
WeightsOpenOpen
Context window131K131K
Max output8K8K
Input $ / M tokens$0.05$0.08
Output $ / M tokens$0.15$0.16
Results tracked2443

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Gemma 3 12B leads

Gemma 3 12B: 31.7 (#280), Gemma 3 27B: 22.5 (#334)

Coding benchmarks
BenchmarkGemma 3 12BGemma 3 27B
SciCode17.4%21.2%
LMArena Coding12811322
Aider Polyglot—4.9%
LiveBench Coding—39.9%

Agentic & Tool Use Too close to call

Gemma 3 12B: 25.5 (#108), Gemma 3 27B: 25.1 (#110)

Agentic & Tool Use benchmarks
BenchmarkGemma 3 12BGemma 3 27B
Berkeley Function Calling Leaderboard30.4%29.5%

Reasoning Gemma 3 27B leads

Gemma 3 12B: 15.7 (#313), Gemma 3 27B: 16.7 (#301)

Reasoning benchmarks
BenchmarkGemma 3 12BGemma 3 27B
CritPt0%0%
Chess Puzzles0%0%
LMArena Hard Prompts13091340
DTBench48.8%52.5%
LMCA4.5%12.3%
Epoch Capabilities Index123.5130.04
Kagi LLM Benchmark—40.4%
LiveBench Reasoning—43.8%
LiveBench Data Analysis—51.5%
LiveBench—50%

Math Gemma 3 27B leads

Gemma 3 12B: 22.3 (#279), Gemma 3 27B: 25.9 (#265)

Math benchmarks
BenchmarkGemma 3 12BGemma 3 27B
OTIS Mock AIME 2024-202516.7%22.5%
LMArena Math13071312
LiveBench Math—55.4%
MATH Level 5—74%

Knowledge Gemma 3 12B leads

Gemma 3 12B: 26.5 (#257), Gemma 3 27B: 25.5 (#261)

Knowledge benchmarks
BenchmarkGemma 3 12BGemma 3 27B
GPQA Diamond39.5%47.7%
Vectara Hallucination Rate4.4%7.4%
LMArena Expert12481304
Confabulations—40.3%

Multimodal Not comparable

Gemma 3 12B: —, Gemma 3 27B: 32.6 (#100)

Multimodal benchmarks
BenchmarkGemma 3 12BGemma 3 27B
LMArena Vision—1164
GeoBench—52%
MindCube46.7%—

Multilingual Gemma 3 27B leads

Gemma 3 12B: 45.7 (#165), Gemma 3 27B: 46.9 (#155)

Multilingual benchmarks
BenchmarkGemma 3 12BGemma 3 27B
LMArena Non-English13181334
LMArena German13701362
LMArena Russian13351349
LMArena Chinese—1346
LMArena French—1368
LMArena Japanese—1287
LMArena Korean—1308
LMArena Spanish—1349

Instruction Following Gemma 3 27B leads

Gemma 3 12B: 68.6 (#186), Gemma 3 27B: 70.6 (#160)

Instruction Following benchmarks
BenchmarkGemma 3 12BGemma 3 27B
LMArena Instruction Following12991321
LiveBench Instruction Following—74.9%

Long Context Gemma 3 12B leads

Gemma 3 12B: 40.0 (#162), Gemma 3 27B: 27.6 (#293)

Long Context benchmarks
BenchmarkGemma 3 12BGemma 3 27B
LMArena Longer Query13171333
Fiction.LiveBench—33.3%

Writing & Preference Gemma 3 27B leads

Gemma 3 12B: 47.5 (#209), Gemma 3 27B: 52.5 (#168)

Writing & Preference benchmarks
BenchmarkGemma 3 12BGemma 3 27B
LMArena Text13341358
LMArena Creative Writing13311346
EQ-Bench Creative Writing11261266
LMArena Multi-Turn13341345
Short-Story Creative Writing—79.9%
LiveBench Language—34.6%

Frequently asked questions

Is Gemma 3 12B better than Gemma 3 27B?

Gemma 3 12B is the stronger model overall, scoring 32.1 to 30.8 on the Noometry Index.

Which is cheaper, Gemma 3 12B or Gemma 3 27B?

Gemma 3 12B is cheaper. It lists at $0.05 per million input tokens and $0.15 per million output tokens; Gemma 3 27B lists at $0.08 and $0.16.

Is Gemma 3 12B or Gemma 3 27B better for coding?

Gemma 3 12B scores higher on coding benchmarks: 31.7 versus 22.5 in the Noometry coding category.

Which has the bigger context window?

Both accept 131K tokens.

How many benchmarks do Gemma 3 12B and Gemma 3 27B share?

23 benchmarks have published results for both models. Gemma 3 12B has 24 scored results on Noometry and Gemma 3 27B has 43.

Related comparisons

Go deeper