Model comparison

Gemma 2 9B vs Gemma 2B

Gemma 2B is the stronger model overall, scoring 29.6 to 25.9 on the Noometry Index.

Last verified . 16 shared benchmarks.

Gemma 2 9B Google

25.9

Rank #341 Confirmed

Gemma 2B Google

29.6

Rank #307 Confirmed

Summary

  • They share 16 benchmarks with published results for both. Gemma 2 9B scores higher in 5 categories and Gemma 2B in 2 categories; 6 gaps are clear of the uncertainty.
  • The widest gap is in math, where Gemma 2B leads 30.0 to 9.9.

Side by side

Gemma 2 9B and Gemma 2B specifications
Gemma 2 9BGemma 2B
ProviderGoogleGoogle
Noometry Index25.929.6
Released2024-06-242024-02-21
WeightsOpenOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked3523

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

Gemma 2 9B: 29.4 (#304), Gemma 2B: 29.4 (#305)

Coding benchmarks
BenchmarkGemma 2 9BGemma 2B
LMArena Coding11731010
BigCodeBench Instruct34.7%—
LiveBench Coding22.5%—
BigCodeBench Complete40.6%—
HumanEval+—20.7%
MBPP+—34.1%

Reasoning Gemma 2B leads

Gemma 2 9B: 15.9 (#309), Gemma 2B: 18.8 (#275)

Reasoning benchmarks
BenchmarkGemma 2 9BGemma 2B
LMArena Hard Prompts1171989
Epoch Capabilities Index119.8394.2
PIQA83.7%77.3%
LiveBench Reasoning15.2%—
LiveBench Data Analysis36.4%—
BIG-Bench Hard—35.2%
HellaSwag—71.4%
LiveBench28.7%—
WinoGrande—65.4%

Math Gemma 2B leads

Gemma 2 9B: 9.9 (#318), Gemma 2B: 30.0 (#239)

Math benchmarks
BenchmarkGemma 2 9BGemma 2B
LMArena Math11831009
GSM8K84.9%17.7%
OTIS Mock AIME 2024-20250.6%—
LiveBench Math19.8%—
MATH Level 521%—

Knowledge Not comparable

Gemma 2 9B: 9.7 (#305), Gemma 2B: —

Knowledge benchmarks
BenchmarkGemma 2 9BGemma 2B
BoolQ85.7%69.4%
MMLU72.1%42.3%
GPQA Diamond27.5%—
LMArena Expert1147—
ARC (AI2) Challenge—42.1%
TriviaQA—53.2%

Multilingual Gemma 2 9B leads

Gemma 2 9B: 36.6 (#238), Gemma 2B: 23.0 (#294)

Multilingual benchmarks
BenchmarkGemma 2 9BGemma 2B
LMArena Non-English1188958
LMArena Chinese1185986
LMArena Russian1200937
LMArena French1190—
LMArena German1186—
LMArena Japanese1144—
LMArena Korean1137—
LMArena Spanish1200—

Instruction Following Gemma 2 9B leads

Gemma 2 9B: 57.6 (#269), Gemma 2B: 48.5 (#302)

Instruction Following benchmarks
BenchmarkGemma 2 9BGemma 2B
LMArena Instruction Following1178970
LiveBench Instruction Following52.6%—

Long Context Gemma 2 9B leads

Gemma 2 9B: 36.3 (#233), Gemma 2B: 29.9 (#291)

Long Context benchmarks
BenchmarkGemma 2 9BGemma 2B
LMArena Longer Query1197981

Writing & Preference Gemma 2 9B leads

Gemma 2 9B: 32.1 (#281), Gemma 2B: 24.0 (#308)

Writing & Preference benchmarks
BenchmarkGemma 2 9BGemma 2B
LMArena Text12071002
LMArena Creative Writing1206987
LMArena Multi-Turn1193945
EQ-Bench Creative Writing841—
LiveBench Language25.5%—

Frequently asked questions

Is Gemma 2 9B better than Gemma 2B?

Gemma 2B is the stronger model overall, scoring 29.6 to 25.9 on the Noometry Index.

Is Gemma 2 9B or Gemma 2B better for coding?

They score almost the same on coding (29.4 vs 29.4); test both on your own repository before choosing.

How many benchmarks do Gemma 2 9B and Gemma 2B share?

16 benchmarks have published results for both models. Gemma 2 9B has 35 scored results on Noometry and Gemma 2B has 23.

Related comparisons

Go deeper