Model comparison

Devstral Small 2505 vs Magistral Medium

Devstral Small 2505 and Magistral Medium score almost the same on the Noometry Index (34.3 vs 35.2), so choose on price, context window or the category you care about most.

Last verified . 3 shared benchmarks.

Devstral Small 2505 Mistral AI

34.3

Rank #233 Reported

Magistral Medium Mistral AI

35.2

Rank #227 Confirmed

Summary

  • They share 3 benchmarks with published results for both. Devstral Small 2505 scores higher in 1 category and Magistral Medium in 1 category; one gap is clear of the uncertainty.
  • The widest gap is in reasoning, where Devstral Small 2505 leads 19.7 to 8.6.
  • The biggest single-benchmark swing is Kagi LLM Benchmark: 37.7% for Devstral Small 2505 and 16.2% for Magistral Medium.
  • Devstral Small 2505 is cheaper at $0.10 / $0.30 per million input/output tokens, against $2 / $5 for Magistral Medium.
  • Magistral Medium accepts more context: 262K tokens versus 128K.

Side by side

Devstral Small 2505 and Magistral Medium specifications
Devstral Small 2505Magistral Medium
ProviderMistral AIMistral AI
Noometry Index34.335.2
Released2025-05-072025-03-17
WeightsOpenOpen
Context window128K262K
Max output128K16K
Input $ / M tokens$0.10$2
Output $ / M tokens$0.30$5
Results tracked422

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

Devstral Small 2505: 38.9 (#166), Magistral Medium: 39.1 (#161)

Coding benchmarks
BenchmarkDevstral Small 2505Magistral Medium
SciCode28.8%39.2%
SWE-bench Verified (bash only)56.4%—
LMArena Coding—1319

Reasoning Devstral Small 2505 leads

Devstral Small 2505: 19.7 (#252), Magistral Medium: 8.6 (#348)

Reasoning benchmarks
BenchmarkDevstral Small 2505Magistral Medium
Kagi LLM Benchmark37.7%16.2%
CritPt0%0.3%
ARC-AGI-2—0%
ARC-AGI-1—6.1%
LMArena Hard Prompts—1267

Math Not comparable

Devstral Small 2505: —, Magistral Medium: 35.1 (#189)

Math benchmarks
BenchmarkDevstral Small 2505Magistral Medium
LMArena Math—1250

Knowledge Not comparable

Devstral Small 2505: —, Magistral Medium: 33.5 (#202)

Knowledge benchmarks
BenchmarkDevstral Small 2505Magistral Medium
LMArena Expert—1223

Multilingual Not comparable

Devstral Small 2505: —, Magistral Medium: 39.6 (#224)

Multilingual benchmarks
BenchmarkDevstral Small 2505Magistral Medium
LMArena Non-English—1232
LMArena Chinese—1227
LMArena French—1267
LMArena German—1248
LMArena Japanese—1175
LMArena Korean—1125
LMArena Russian—1224
LMArena Spanish—1271

Instruction Following Not comparable

Devstral Small 2505: —, Magistral Medium: 66.0 (#211)

Instruction Following benchmarks
BenchmarkDevstral Small 2505Magistral Medium
LMArena Instruction Following—1254

Long Context Not comparable

Devstral Small 2505: —, Magistral Medium: 39.3 (#183)

Long Context benchmarks
BenchmarkDevstral Small 2505Magistral Medium
LMArena Longer Query—1295

Writing & Preference Not comparable

Devstral Small 2505: —, Magistral Medium: 46.3 (#219)

Writing & Preference benchmarks
BenchmarkDevstral Small 2505Magistral Medium
LMArena Text—1255
LMArena Creative Writing—1245
LMArena Multi-Turn—1275

Frequently asked questions

Is Devstral Small 2505 better than Magistral Medium?

Devstral Small 2505 and Magistral Medium score almost the same on the Noometry Index (34.3 vs 35.2), so choose on price, context window or the category you care about most.

Which is cheaper, Devstral Small 2505 or Magistral Medium?

Devstral Small 2505 is cheaper. It lists at $0.10 per million input tokens and $0.30 per million output tokens; Magistral Medium lists at $2 and $5.

Is Devstral Small 2505 or Magistral Medium better for coding?

They score almost the same on coding (38.9 vs 39.1); test both on your own repository before choosing.

Which has the bigger context window?

Magistral Medium does, with 262K tokens against 128K.

How many benchmarks do Devstral Small 2505 and Magistral Medium share?

3 benchmarks have published results for both models. Devstral Small 2505 has 4 scored results on Noometry and Magistral Medium has 22.

Related comparisons

Go deeper