Model comparison

Devstral Small 2505 vs Gemma 7B

Devstral Small 2505 is the stronger model overall, scoring 34.3 to 30.0 on the Noometry Index.

Last verified . 0 shared benchmarks.

Devstral Small 2505 Mistral AI

34.3

Rank #233 Reported

Gemma 7B Google

30.0

Rank #299 Confirmed

Summary

  • The widest gap is in coding, where Devstral Small 2505 leads 38.9 to 30.5.

Side by side

Devstral Small 2505 and Gemma 7B specifications
Devstral Small 2505Gemma 7B
ProviderMistral AIGoogle
Noometry Index34.330.0
Released2025-05-072024-02-21
WeightsOpenOpen
Context window128K—
Max output128K—
Input $ / M tokens$0.10—
Output $ / M tokens$0.30—
Results tracked427

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Devstral Small 2505 leads

Devstral Small 2505: 38.9 (#166), Gemma 7B: 30.5 (#294)

Coding benchmarks
BenchmarkDevstral Small 2505Gemma 7B
SWE-bench Verified (bash only)56.4%—
SciCode28.8%—
LMArena Coding—1048
HumanEval+—28.7%
MBPP+—43.4%

Reasoning Too close to call

Devstral Small 2505: 19.7 (#252), Gemma 7B: 19.9 (#249)

Reasoning benchmarks
BenchmarkDevstral Small 2505Gemma 7B
Kagi LLM Benchmark37.7%—
CritPt0%—
LMArena Hard Prompts—1042
Adversarial NLI—48.7%
BIG-Bench Hard—55.1%
Epoch Capabilities Index—111.99
HellaSwag—82.2%
PIQA—81.2%
WinoGrande—79%

Math Not comparable

Devstral Small 2505: —, Gemma 7B: 31.2 (#228)

Math benchmarks
BenchmarkDevstral Small 2505Gemma 7B
LMArena Math—1066
GSM8K—46.4%

Knowledge Not comparable

Devstral Small 2505: —, Gemma 7B: 27.3 (#252)

Knowledge benchmarks
BenchmarkDevstral Small 2505Gemma 7B
LMArena Expert—1001
ARC (AI2) Challenge—78.3%
BoolQ—83.2%
MMLU—66.1%
OpenBookQA—78.6%
TriviaQA—72.3%

Multilingual Not comparable

Devstral Small 2505: —, Gemma 7B: 25.1 (#287)

Multilingual benchmarks
BenchmarkDevstral Small 2505Gemma 7B
LMArena Non-English—999
LMArena Chinese—1035
LMArena French—1025
LMArena Russian—993

Instruction Following Not comparable

Devstral Small 2505: —, Gemma 7B: 51.5 (#295)

Instruction Following benchmarks
BenchmarkDevstral Small 2505Gemma 7B
LMArena Instruction Following—1017

Long Context Not comparable

Devstral Small 2505: —, Gemma 7B: 31.1 (#282)

Long Context benchmarks
BenchmarkDevstral Small 2505Gemma 7B
LMArena Longer Query—1022

Writing & Preference Not comparable

Devstral Small 2505: —, Gemma 7B: 27.1 (#302)

Writing & Preference benchmarks
BenchmarkDevstral Small 2505Gemma 7B
LMArena Text—1056
LMArena Creative Writing—1024
LMArena Multi-Turn—963

Frequently asked questions

Is Devstral Small 2505 better than Gemma 7B?

Devstral Small 2505 is the stronger model overall, scoring 34.3 to 30.0 on the Noometry Index.

Is Devstral Small 2505 or Gemma 7B better for coding?

Devstral Small 2505 scores higher on coding benchmarks: 38.9 versus 30.5 in the Noometry coding category.

How many benchmarks do Devstral Small 2505 and Gemma 7B share?

0 benchmarks have published results for both models. Devstral Small 2505 has 4 scored results on Noometry and Gemma 7B has 27.

Related comparisons

Go deeper