Model comparison

Devstral Small 2505 vs o1-mini

Devstral Small 2505 and o1-mini score almost the same on the Noometry Index (34.3 vs 34.0), so choose on price, context window or the category you care about most.

Last verified . 0 shared benchmarks.

Devstral Small 2505 Mistral AI

34.3

Rank #233 Reported

o1-mini OpenAI

34.0

Rank #235 Confirmed

Summary

  • The widest gap is in reasoning, where Devstral Small 2505 leads 19.7 to 8.8.
  • Devstral Small 2505 has downloadable open weights; the other is API-only.

Side by side

Devstral Small 2505 and o1-mini specifications
Devstral Small 2505o1-mini
ProviderMistral AIOpenAI
Noometry Index34.334.0
Released2025-05-072024-09-12
WeightsOpenProprietary
Context window128K—
Max output128K—
Input $ / M tokens$0.10—
Output $ / M tokens$0.30—
Results tracked439

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Devstral Small 2505 leads

Devstral Small 2505: 38.9 (#166), o1-mini: 35.5 (#224)

Coding benchmarks
BenchmarkDevstral Small 2505o1-mini
SWE-bench Verified (bash only)56.4%—
Aider Polyglot—32.9%
SciCode28.8%—
WeirdML—36.3%
LiveBench Coding—48%
LMArena Coding—1362
HumanEval+—89%
MBPP+—78.8%

Agentic & Tool Use Not comparable

Devstral Small 2505: —, o1-mini: 24.6 (#118)

Agentic & Tool Use benchmarks
BenchmarkDevstral Small 2505o1-mini
Cybench—10%

Reasoning Devstral Small 2505 leads

Devstral Small 2505: 19.7 (#252), o1-mini: 8.8 (#346)

Reasoning benchmarks
BenchmarkDevstral Small 2505o1-mini
ARC-AGI-2—0.8%
SimpleBench—18.1%
Kagi LLM Benchmark37.7%—
ARC-AGI-1—14%
CritPt0%—
LiveBench Reasoning—72.3%
LMArena Hard Prompts—1333
LiveBench Data Analysis—57.9%
Epoch Capabilities Index—135.82
LiveBench—57.8%

Math Not comparable

Devstral Small 2505: —, o1-mini: 35.4 (#186)

Math benchmarks
BenchmarkDevstral Small 2505o1-mini
OTIS Mock AIME 2024-2025—46.9%
LiveBench Math—62%
LMArena Math—1358
MATH Level 5—89.2%
FrontierMath (Feb 2025 set)—1.7%

Knowledge Not comparable

Devstral Small 2505: —, o1-mini: 34.9 (#192)

Knowledge benchmarks
BenchmarkDevstral Small 2505o1-mini
GPQA Diamond—62.4%
Confabulations—18.6%
LMArena Expert—1316

Multilingual Not comparable

Devstral Small 2505: —, o1-mini: 43.6 (#182)

Multilingual benchmarks
BenchmarkDevstral Small 2505o1-mini
LMArena Non-English—1289
LMArena Chinese—1314
LMArena French—1293
LMArena German—1278
LMArena Japanese—1245
LMArena Korean—1223
LMArena Russian—1283
LMArena Spanish—1303

Instruction Following Not comparable

Devstral Small 2505: —, o1-mini: 66.7 (#206)

Instruction Following benchmarks
BenchmarkDevstral Small 2505o1-mini
LiveBench Instruction Following—65.4%
LMArena Instruction Following—1304

Long Context Not comparable

Devstral Small 2505: —, o1-mini: 40.1 (#161)

Long Context benchmarks
BenchmarkDevstral Small 2505o1-mini
LMArena Longer Query—1320

Writing & Preference Not comparable

Devstral Small 2505: —, o1-mini: 48.4 (#202)

Writing & Preference benchmarks
BenchmarkDevstral Small 2505o1-mini
LMArena Text—1317
LMArena Creative Writing—1244
Short-Story Creative Writing—64.9%
LMArena Multi-Turn—1314
LiveBench Language—40.9%

Frequently asked questions

Is Devstral Small 2505 better than o1-mini?

Devstral Small 2505 and o1-mini score almost the same on the Noometry Index (34.3 vs 34.0), so choose on price, context window or the category you care about most.

Is Devstral Small 2505 or o1-mini better for coding?

Devstral Small 2505 scores higher on coding benchmarks: 38.9 versus 35.5 in the Noometry coding category.

How many benchmarks do Devstral Small 2505 and o1-mini share?

0 benchmarks have published results for both models. Devstral Small 2505 has 4 scored results on Noometry and o1-mini has 39.

Related comparisons

Go deeper