Model comparison

Claude 3.5 Sonnet vs Magistral Medium

Claude 3.5 Sonnet and Magistral Medium score almost the same on the Noometry Index (34.6 vs 35.2), so choose on price, context window or the category you care about most.

Last verified . 17 shared benchmarks.

Claude 3.5 Sonnet Anthropic

34.6

Rank #231 Confirmed

Magistral Medium Mistral AI

35.2

Rank #227 Confirmed

Summary

  • They share 17 benchmarks with published results for both. Claude 3.5 Sonnet scores higher in 5 categories and Magistral Medium in 3 categories; 6 gaps are clear of the uncertainty.
  • The widest gap is in math, where Magistral Medium leads 35.1 to 19.2.
  • Magistral Medium has downloadable open weights; the other is API-only.

Side by side

Claude 3.5 Sonnet and Magistral Medium specifications
Claude 3.5 SonnetMagistral Medium
ProviderAnthropicMistral AI
Noometry Index34.635.2
Released2024-06-202025-03-17
WeightsProprietaryOpen
Context window—262K
Max output—16K
Input $ / M tokens—$2
Output $ / M tokens—$5
Results tracked6022

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

Claude 3.5 Sonnet: 39.0 (#165), Magistral Medium: 39.1 (#161)

Coding benchmarks
BenchmarkClaude 3.5 SonnetMagistral Medium
LMArena Coding13421319
Aider Polyglot51.6%—
SciCode—39.2%
GSO4.6%—
WeirdML40%—
BigCodeBench Instruct46.8%—
LiveBench Coding67.1%—
BigCodeBench Complete58.6%—
CadEval48%—
HumanEval+81.7%—
MBPP+74.3%—

Agentic & Tool Use Not comparable

Claude 3.5 Sonnet: 32.3 (#67), Magistral Medium: —

Agentic & Tool Use benchmarks
BenchmarkClaude 3.5 SonnetMagistral Medium
TheAgentCompany24%—
Cybench17.5%—
BALROG32.6%—
METR Time Horizons45.2%—

Reasoning Claude 3.5 Sonnet leads

Claude 3.5 Sonnet: 23.1 (#183), Magistral Medium: 8.6 (#348)

Reasoning benchmarks
BenchmarkClaude 3.5 SonnetMagistral Medium
LMArena Hard Prompts13051267
ARC-AGI-2—0%
SimpleBench41.4%—
Kagi LLM Benchmark—16.2%
ARC-AGI-1—6.1%
CritPt—0.3%
EnigmaEval0.9%—
LiveBench Reasoning56.7%—
DTBench67.8%—
LiveBench Data Analysis55%—
Epoch Capabilities Index133.55—
ForecastBench60.7—
LiveBench59%—

Math Magistral Medium leads

Claude 3.5 Sonnet: 19.2 (#288), Magistral Medium: 35.1 (#189)

Math benchmarks
BenchmarkClaude 3.5 SonnetMagistral Medium
LMArena Math13071250
OTIS Mock AIME 2024-20258.5%—
Omni-MATH27.6%—
LiveBench Math52.3%—
MATH Level 556.9%—
FrontierMath (Feb 2025 set)2.1%—
FrontierMath Tier 4 (v1)0%—

Knowledge Magistral Medium leads

Claude 3.5 Sonnet: 28.6 (#245), Magistral Medium: 33.5 (#202)

Knowledge benchmarks
BenchmarkClaude 3.5 SonnetMagistral Medium
LMArena Expert12651223
GPQA Diamond55.3%—
Humanity's Last Exam4.1%—
MMLU-Pro77.7%—
Confabulations19.9%—
GPQA (HELM)56.5%—
MMLU87.3%—

Multimodal Not comparable

Claude 3.5 Sonnet: 26.5 (#120), Magistral Medium: —

Multimodal benchmarks
BenchmarkClaude 3.5 SonnetMagistral Medium
LMArena Vision1125—
Video-MME60%—
GeoBench62%—
VPCT33%—

Multilingual Claude 3.5 Sonnet leads

Claude 3.5 Sonnet: 43.2 (#185), Magistral Medium: 39.6 (#224)

Multilingual benchmarks
BenchmarkClaude 3.5 SonnetMagistral Medium
LMArena Non-English12831232
LMArena Chinese12721227
LMArena French13051267
LMArena German12971248
LMArena Japanese12341175
LMArena Korean12001125
LMArena Russian13061224
LMArena Spanish12901271

Instruction Following Claude 3.5 Sonnet leads

Claude 3.5 Sonnet: 68.8 (#182), Magistral Medium: 66.0 (#211)

Instruction Following benchmarks
BenchmarkClaude 3.5 SonnetMagistral Medium
LMArena Instruction Following12971254
LiveBench Instruction Following69.3%—
IFEval85.5%—

Long Context Too close to call

Claude 3.5 Sonnet: 39.9 (#167), Magistral Medium: 39.3 (#183)

Long Context benchmarks
BenchmarkClaude 3.5 SonnetMagistral Medium
LMArena Longer Query13111295

Writing & Preference Claude 3.5 Sonnet leads

Claude 3.5 Sonnet: 52.9 (#164), Magistral Medium: 46.3 (#219)

Writing & Preference benchmarks
BenchmarkClaude 3.5 SonnetMagistral Medium
LMArena Text12981255
LMArena Creative Writing12921245
LMArena Multi-Turn13261275
Short-Story Creative Writing80.3%—
EQ-Bench Creative Writing1451—
WildBench79.2%—
LiveBench Language53.8%—

Frequently asked questions

Is Claude 3.5 Sonnet better than Magistral Medium?

Claude 3.5 Sonnet and Magistral Medium score almost the same on the Noometry Index (34.6 vs 35.2), so choose on price, context window or the category you care about most.

Is Claude 3.5 Sonnet or Magistral Medium better for coding?

They score almost the same on coding (39.0 vs 39.1); test both on your own repository before choosing.

How many benchmarks do Claude 3.5 Sonnet and Magistral Medium share?

17 benchmarks have published results for both models. Claude 3.5 Sonnet has 60 scored results on Noometry and Magistral Medium has 22.

Related comparisons

Go deeper