Model comparison

DeepSeek-R1-Distill-Qwen-32B vs Magistral Medium

DeepSeek-R1-Distill-Qwen-32B and Magistral Medium score almost the same on the Noometry Index (35.5 vs 35.2), so choose on price, context window or the category you care about most.

Last verified . 0 shared benchmarks.

DeepSeek-R1-Distill-Qwen-32B DeepSeek

35.5

Rank #226 Confirmed

Magistral Medium Mistral AI

35.2

Rank #227 Confirmed

Summary

  • The widest gap is in reasoning, where DeepSeek-R1-Distill-Qwen-32B leads 18.2 to 8.6.

Side by side

DeepSeek-R1-Distill-Qwen-32B and Magistral Medium specifications
DeepSeek-R1-Distill-Qwen-32BMagistral Medium
ProviderDeepSeekMistral AI
Noometry Index35.535.2
Released2025-01-202025-03-17
WeightsOpenOpen
Context window—262K
Max output—16K
Input $ / M tokens—$2
Output $ / M tokens—$5
Results tracked1422

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Magistral Medium leads

DeepSeek-R1-Distill-Qwen-32B: 36.1 (#212), Magistral Medium: 39.1 (#161)

Coding benchmarks
BenchmarkDeepSeek-R1-Distill-Qwen-32BMagistral Medium
SciCode—39.2%
BigCodeBench Instruct43.9%—
LiveBench Coding33.7%—
LMArena Coding—1319
BigCodeBench Complete54.9%—

Agentic & Tool Use Not comparable

DeepSeek-R1-Distill-Qwen-32B: 28.1 (#94), Magistral Medium: —

Agentic & Tool Use benchmarks
BenchmarkDeepSeek-R1-Distill-Qwen-32BMagistral Medium
BALROG19.5%—

Reasoning DeepSeek-R1-Distill-Qwen-32B leads

DeepSeek-R1-Distill-Qwen-32B: 18.2 (#284), Magistral Medium: 8.6 (#348)

Reasoning benchmarks
BenchmarkDeepSeek-R1-Distill-Qwen-32BMagistral Medium
ARC-AGI-2—0%
Kagi LLM Benchmark—16.2%
ARC-AGI-1—6.1%
CritPt—0.3%
Chess Puzzles1%—
LiveBench Reasoning52.3%—
LMArena Hard Prompts—1267
LiveBench Data Analysis45.4%—
Epoch Capabilities Index137.44—
LiveBench45.5%—

Math Too close to call

DeepSeek-R1-Distill-Qwen-32B: 34.5 (#194), Magistral Medium: 35.1 (#189)

Math benchmarks
BenchmarkDeepSeek-R1-Distill-Qwen-32BMagistral Medium
OTIS Mock AIME 2024-202555.6%—
LiveBench Math59.4%—
LMArena Math—1250

Knowledge DeepSeek-R1-Distill-Qwen-32B leads

DeepSeek-R1-Distill-Qwen-32B: 35.7 (#182), Magistral Medium: 33.5 (#202)

Knowledge benchmarks
BenchmarkDeepSeek-R1-Distill-Qwen-32BMagistral Medium
GPQA Diamond64.1%—
LMArena Expert—1223

Multilingual Not comparable

DeepSeek-R1-Distill-Qwen-32B: —, Magistral Medium: 39.6 (#224)

Multilingual benchmarks
BenchmarkDeepSeek-R1-Distill-Qwen-32BMagistral Medium
LMArena Non-English—1232
LMArena Chinese—1227
LMArena French—1267
LMArena German—1248
LMArena Japanese—1175
LMArena Korean—1125
LMArena Russian—1224
LMArena Spanish—1271

Instruction Following Magistral Medium leads

DeepSeek-R1-Distill-Qwen-32B: 61.6 (#243), Magistral Medium: 66.0 (#211)

Instruction Following benchmarks
BenchmarkDeepSeek-R1-Distill-Qwen-32BMagistral Medium
LiveBench Instruction Following55.7%—
LMArena Instruction Following—1254

Long Context Not comparable

DeepSeek-R1-Distill-Qwen-32B: —, Magistral Medium: 39.3 (#183)

Long Context benchmarks
BenchmarkDeepSeek-R1-Distill-Qwen-32BMagistral Medium
LMArena Longer Query—1295

Writing & Preference DeepSeek-R1-Distill-Qwen-32B leads

DeepSeek-R1-Distill-Qwen-32B: 49.6 (#188), Magistral Medium: 46.3 (#219)

Writing & Preference benchmarks
BenchmarkDeepSeek-R1-Distill-Qwen-32BMagistral Medium
LMArena Text—1255
LMArena Creative Writing—1245
LMArena Multi-Turn—1275
LiveBench Language26.8%—

Frequently asked questions

Is DeepSeek-R1-Distill-Qwen-32B better than Magistral Medium?

DeepSeek-R1-Distill-Qwen-32B and Magistral Medium score almost the same on the Noometry Index (35.5 vs 35.2), so choose on price, context window or the category you care about most.

Is DeepSeek-R1-Distill-Qwen-32B or Magistral Medium better for coding?

Magistral Medium scores higher on coding benchmarks: 39.1 versus 36.1 in the Noometry coding category.

How many benchmarks do DeepSeek-R1-Distill-Qwen-32B and Magistral Medium share?

0 benchmarks have published results for both models. DeepSeek-R1-Distill-Qwen-32B has 14 scored results on Noometry and Magistral Medium has 22.

Related comparisons

Go deeper