Model comparison

Devstral Small 2505 vs Grok-2 (Dec 2024)

Devstral Small 2505 and Grok-2 (Dec 2024) score almost the same on the Noometry Index (34.3 vs 33.7), so choose on price, context window or the category you care about most.

Last verified . 0 shared benchmarks.

Devstral Small 2505 Mistral AI

34.3

Rank #233 Reported

Grok-2 (Dec 2024) xAI

33.7

Rank #239 Confirmed

Summary

  • The widest gap is in coding, where Devstral Small 2505 leads 38.9 to 33.3.
  • Devstral Small 2505 has downloadable open weights; the other is API-only.

Side by side

Devstral Small 2505 and Grok-2 (Dec 2024) specifications
Devstral Small 2505Grok-2 (Dec 2024)
ProviderMistral AIxAI
Noometry Index34.333.7
Released2025-05-072024-08-13
WeightsOpenProprietary
Context window128K—
Max output128K—
Input $ / M tokens$0.10—
Output $ / M tokens$0.30—
Results tracked434

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Devstral Small 2505 leads

Devstral Small 2505: 38.9 (#166), Grok-2 (Dec 2024): 33.3 (#258)

Coding benchmarks
BenchmarkDevstral Small 2505Grok-2 (Dec 2024)
SWE-bench Verified (bash only)56.4%—
SciCode28.8%—
WeirdML—22.2%
LiveBench Coding—46.4%
LMArena Coding—1287

Reasoning Devstral Small 2505 leads

Devstral Small 2505: 19.7 (#252), Grok-2 (Dec 2024): 16.9 (#299)

Reasoning benchmarks
BenchmarkDevstral Small 2505Grok-2 (Dec 2024)
SimpleBench—22.7%
Kagi LLM Benchmark37.7%—
CritPt0%—
LiveBench Reasoning—54.8%
LMArena Hard Prompts—1272
DTBench—65.2%
LiveBench Data Analysis—54.5%
Epoch Capabilities Index—130.48
LiveBench—54.3%

Math Not comparable

Devstral Small 2505: —, Grok-2 (Dec 2024): 20.8 (#284)

Math benchmarks
BenchmarkDevstral Small 2505Grok-2 (Dec 2024)
OTIS Mock AIME 2024-2025—11.5%
LiveBench Math—54.9%
LMArena Math—1283
MATH Level 5—63.5%
FrontierMath (Feb 2025 set)—0.7%

Knowledge Not comparable

Devstral Small 2505: —, Grok-2 (Dec 2024): 29.8 (#233)

Knowledge benchmarks
BenchmarkDevstral Small 2505Grok-2 (Dec 2024)
GPQA Diamond—53.8%
Confabulations—20.1%
LMArena Expert—1254

Multilingual Not comparable

Devstral Small 2505: —, Grok-2 (Dec 2024): 43.1 (#188)

Multilingual benchmarks
BenchmarkDevstral Small 2505Grok-2 (Dec 2024)
LMArena Non-English—1282
LMArena Chinese—1289
LMArena French—1318
LMArena German—1287
LMArena Japanese—1244
LMArena Korean—1237
LMArena Russian—1286
LMArena Spanish—1281

Instruction Following Not comparable

Devstral Small 2505: —, Grok-2 (Dec 2024): 66.9 (#202)

Instruction Following benchmarks
BenchmarkDevstral Small 2505Grok-2 (Dec 2024)
LiveBench Instruction Following—69.6%
LMArena Instruction Following—1270

Long Context Not comparable

Devstral Small 2505: —, Grok-2 (Dec 2024): 38.8 (#190)

Long Context benchmarks
BenchmarkDevstral Small 2505Grok-2 (Dec 2024)
LMArena Longer Query—1276

Writing & Preference Not comparable

Devstral Small 2505: —, Grok-2 (Dec 2024): 48.6 (#198)

Writing & Preference benchmarks
BenchmarkDevstral Small 2505Grok-2 (Dec 2024)
LMArena Text—1305
LMArena Creative Writing—1284
Short-Story Creative Writing—63.6%
LMArena Multi-Turn—1290
LiveBench Language—45.6%

Frequently asked questions

Is Devstral Small 2505 better than Grok-2 (Dec 2024)?

Devstral Small 2505 and Grok-2 (Dec 2024) score almost the same on the Noometry Index (34.3 vs 33.7), so choose on price, context window or the category you care about most.

Is Devstral Small 2505 or Grok-2 (Dec 2024) better for coding?

Devstral Small 2505 scores higher on coding benchmarks: 38.9 versus 33.3 in the Noometry coding category.

How many benchmarks do Devstral Small 2505 and Grok-2 (Dec 2024) share?

0 benchmarks have published results for both models. Devstral Small 2505 has 4 scored results on Noometry and Grok-2 (Dec 2024) has 34.

Related comparisons

Go deeper