Model comparison

Dolly 2.0-12b vs Gemma 2 9B

Dolly 2.0-12b and Gemma 2 9B score almost the same on the Noometry Index (25.5 vs 25.9), so choose on price, context window or the category you care about most.

Last verified . 13 shared benchmarks.

Dolly 2.0-12b Databricks

25.5

Rank #342 Confirmed

Gemma 2 9B Google

25.9

Rank #341 Confirmed

Summary

  • They share 13 benchmarks with published results for both. Dolly 2.0-12b scores higher in 1 category and Gemma 2 9B in 5 categories; 5 gaps are clear of the uncertainty.
  • The widest gap is in multilingual, where Gemma 2 9B leads 36.6 to 17.4.

Side by side

Dolly 2.0-12b and Gemma 2 9B specifications
Dolly 2.0-12bGemma 2 9B
ProviderDatabricksGoogle
Noometry Index25.525.9
Released2023-04-112024-06-24
WeightsOpenOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked1735

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Gemma 2 9B leads

Dolly 2.0-12b: 23.4 (#332), Gemma 2 9B: 29.4 (#304)

Coding benchmarks
BenchmarkDolly 2.0-12bGemma 2 9B
LMArena Coding7761173
BigCodeBench Instruct—34.7%
LiveBench Coding—22.5%
BigCodeBench Complete—40.6%

Reasoning Too close to call

Dolly 2.0-12b: 15.3 (#316), Gemma 2 9B: 15.9 (#309)

Reasoning benchmarks
BenchmarkDolly 2.0-12bGemma 2 9B
LMArena Hard Prompts8041171
Epoch Capabilities Index89.67119.83
PIQA75.4%83.7%
LiveBench Reasoning—15.2%
LiveBench Data Analysis—36.4%
HellaSwag70.8%—
LiveBench—28.7%
WinoGrande61.8%—

Math Dolly 2.0-12b leads

Dolly 2.0-12b: 27.3 (#251), Gemma 2 9B: 9.9 (#318)

Math benchmarks
BenchmarkDolly 2.0-12bGemma 2 9B
LMArena Math8711183
OTIS Mock AIME 2024-2025—0.6%
LiveBench Math—19.8%
MATH Level 5—21%
GSM8K—84.9%

Knowledge Not comparable

Dolly 2.0-12b: —, Gemma 2 9B: 9.7 (#305)

Knowledge benchmarks
BenchmarkDolly 2.0-12bGemma 2 9B
BoolQ56.3%85.7%
MMLU26.2%72.1%
GPQA Diamond—27.5%
LMArena Expert—1147
ARC (AI2) Challenge39.6%—
OpenBookQA39.2%—

Multilingual Gemma 2 9B leads

Dolly 2.0-12b: 17.4 (#296), Gemma 2 9B: 36.6 (#238)

Multilingual benchmarks
BenchmarkDolly 2.0-12bGemma 2 9B
LMArena Non-English8361188
LMArena Chinese8361185
LMArena French—1190
LMArena German—1186
LMArena Japanese—1144
LMArena Korean—1137
LMArena Russian—1200
LMArena Spanish—1200

Instruction Following Gemma 2 9B leads

Dolly 2.0-12b: 38.7 (#304), Gemma 2 9B: 57.6 (#269)

Instruction Following benchmarks
BenchmarkDolly 2.0-12bGemma 2 9B
LMArena Instruction Following8141178
LiveBench Instruction Following—52.6%

Long Context Not comparable

Dolly 2.0-12b: —, Gemma 2 9B: 36.3 (#233)

Long Context benchmarks
BenchmarkDolly 2.0-12bGemma 2 9B
LMArena Longer Query—1197

Writing & Preference Gemma 2 9B leads

Dolly 2.0-12b: 15.2 (#311), Gemma 2 9B: 32.1 (#281)

Writing & Preference benchmarks
BenchmarkDolly 2.0-12bGemma 2 9B
LMArena Text8511207
LMArena Creative Writing8641206
LMArena Multi-Turn7401193
EQ-Bench Creative Writing—841
LiveBench Language—25.5%

Frequently asked questions

Is Dolly 2.0-12b better than Gemma 2 9B?

Dolly 2.0-12b and Gemma 2 9B score almost the same on the Noometry Index (25.5 vs 25.9), so choose on price, context window or the category you care about most.

Is Dolly 2.0-12b or Gemma 2 9B better for coding?

Gemma 2 9B scores higher on coding benchmarks: 29.4 versus 23.4 in the Noometry coding category.

How many benchmarks do Dolly 2.0-12b and Gemma 2 9B share?

13 benchmarks have published results for both models. Dolly 2.0-12b has 17 scored results on Noometry and Gemma 2 9B has 35.

Related comparisons

Go deeper