Model comparison

Gemma 2 9B vs Mistral 7B

Gemma 2 9B is the stronger model overall, scoring 25.9 to 23.0 on the Noometry Index.

Last verified . 26 shared benchmarks.

Gemma 2 9B Google

25.9

Rank #341 Confirmed

Mistral 7B Mistral AI

23.0

Rank #351 Confirmed

Summary

  • They share 26 benchmarks with published results for both. Gemma 2 9B scores higher in 8 categories and Mistral 7B in 0 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in multilingual, where Gemma 2 9B leads 36.6 to 25.8.
  • The biggest single-benchmark swing is MATH Level 5: 21% for Gemma 2 9B and 3.7% for Mistral 7B.

Side by side

Gemma 2 9B and Mistral 7B specifications
Gemma 2 9BMistral 7B
ProviderGoogleMistral AI
Noometry Index25.923.0
Released2024-06-242023-09-27
WeightsOpenOpen
Context window—8K
Max output—8K
Input $ / M tokens—$0.25
Output $ / M tokens—$0.25
Results tracked3537

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Gemma 2 9B leads

Gemma 2 9B: 29.4 (#304), Mistral 7B: 26.4 (#326)

Coding benchmarks
BenchmarkGemma 2 9BMistral 7B
BigCodeBench Instruct34.7%19.5%
LMArena Coding11731082
BigCodeBench Complete40.6%27.3%
LiveBench Coding22.5%—
HumanEval+—36%
MBPP+—42.1%

Reasoning Gemma 2 9B leads

Gemma 2 9B: 15.9 (#309), Mistral 7B: 13.1 (#336)

Reasoning benchmarks
BenchmarkGemma 2 9BMistral 7B
LMArena Hard Prompts11711067
Epoch Capabilities Index119.83112.21
PIQA83.7%83%
Chess Puzzles—0%
LiveBench Reasoning15.2%—
DTBench—42.5%
LiveBench Data Analysis36.4%—
Adversarial NLI—47.1%
BIG-Bench Hard—56.1%
HellaSwag—81%
LiveBench28.7%—
WinoGrande—75.3%

Math Gemma 2 9B leads

Gemma 2 9B: 9.9 (#318), Mistral 7B: 8.1 (#325)

Math benchmarks
BenchmarkGemma 2 9BMistral 7B
OTIS Mock AIME 2024-20250.6%0.3%
LMArena Math11831085
MATH Level 521%3.7%
GSM8K84.9%54.4%
LiveBench Math19.8%—

Knowledge Gemma 2 9B leads

Gemma 2 9B: 9.7 (#305), Mistral 7B: 7.4 (#311)

Knowledge benchmarks
BenchmarkGemma 2 9BMistral 7B
GPQA Diamond27.5%15.2%
LMArena Expert11471036
BoolQ85.7%87.4%
MMLU72.1%62.5%
ARC (AI2) Challenge—78.6%
OpenBookQA—79.8%
TriviaQA—75.2%

Multilingual Gemma 2 9B leads

Gemma 2 9B: 36.6 (#238), Mistral 7B: 25.8 (#283)

Multilingual benchmarks
BenchmarkGemma 2 9BMistral 7B
LMArena Non-English11881012
LMArena Chinese11851009
LMArena French11901037
LMArena German1186987
LMArena Japanese1144878
LMArena Russian12001018
LMArena Spanish12001026
LMArena Korean1137—

Instruction Following Gemma 2 9B leads

Gemma 2 9B: 57.6 (#269), Mistral 7B: 54.2 (#280)

Instruction Following benchmarks
BenchmarkGemma 2 9BMistral 7B
LMArena Instruction Following11781060
LiveBench Instruction Following52.6%—

Long Context Gemma 2 9B leads

Gemma 2 9B: 36.3 (#233), Mistral 7B: 32.2 (#271)

Long Context benchmarks
BenchmarkGemma 2 9BMistral 7B
LMArena Longer Query11971060

Writing & Preference Gemma 2 9B leads

Gemma 2 9B: 32.1 (#281), Mistral 7B: 30.7 (#286)

Writing & Preference benchmarks
BenchmarkGemma 2 9BMistral 7B
LMArena Text12071090
LMArena Creative Writing12061068
LMArena Multi-Turn11931062
EQ-Bench Creative Writing841—
LiveBench Language25.5%—

Frequently asked questions

Is Gemma 2 9B better than Mistral 7B?

Gemma 2 9B is the stronger model overall, scoring 25.9 to 23.0 on the Noometry Index.

Is Gemma 2 9B or Mistral 7B better for coding?

Gemma 2 9B scores higher on coding benchmarks: 29.4 versus 26.4 in the Noometry coding category.

How many benchmarks do Gemma 2 9B and Mistral 7B share?

26 benchmarks have published results for both models. Gemma 2 9B has 35 scored results on Noometry and Mistral 7B has 37.

Related comparisons

Go deeper