Model comparison

Mistral Medium vs o3-mini

Mistral Medium and o3-mini score almost the same on the Noometry Index (36.3 vs 36.7), so choose on price, context window or the category you care about most.

Last verified . 27 shared benchmarks.

Mistral Medium Mistral AI

36.3

Rank #218 Confirmed

o3-mini OpenAI

36.7

Rank #212 Confirmed

Summary

  • They share 27 benchmarks with published results for both. Mistral Medium scores higher in 4 categories and o3-mini in 5 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where o3-mini leads 38.3 to 25.0.
  • The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 32.2% for Mistral Medium and 76.9% for o3-mini.
  • o3-mini is cheaper at $1.10 / $4.40 per million input/output tokens, against $1.50 / $7.50 for Mistral Medium.
  • Mistral Medium accepts more context: 262K tokens versus 200K.
  • Mistral Medium has downloadable open weights; the other is API-only.

Side by side

Mistral Medium and o3-mini specifications
Mistral Mediumo3-mini
ProviderMistral AIOpenAI
Noometry Index36.336.7
Released2023-12-112024-12-20
WeightsOpenProprietary
Context window262K200K
Max output262K100K
Input $ / M tokens$1.50$1.10
Output $ / M tokens$7.50$4.40
Results tracked3651

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding o3-mini leads

Mistral Medium: 34.2 (#243), o3-mini: 40.8 (#132)

Coding benchmarks
BenchmarkMistral Mediumo3-mini
SciCode40.2%39.8%
WeirdML43.7%43.7%
LMArena Coding14341378
FrontierCode8%—
Aider Polyglot—60.4%
GSO—1.3%
LiveBench Coding—82.7%
CadEval—54%
ALE-Bench763.98—

Agentic & Tool Use o3-mini leads

Mistral Medium: 28.3 (#90), o3-mini: 29.6 (#84)

Agentic & Tool Use benchmarks
BenchmarkMistral Mediumo3-mini
Berkeley Function Calling Leaderboard37.7%—
Cybench—22.5%

Reasoning Mistral Medium leads

Mistral Medium: 24.0 (#167), o3-mini: 16.3 (#305)

Reasoning benchmarks
BenchmarkMistral Mediumo3-mini
CritPt0%0.3%
LMArena Hard Prompts14261366
DTBench75.5%68.8%
LMCA26.1%19%
ARC-AGI-2—3%
SimpleBench—22.8%
Kagi LLM Benchmark50%—
ARC-AGI-1—34.5%
Chess Puzzles—17%
LiveBench Reasoning—89.6%
Mystery Game Puzzles—7%
LiveBench Data Analysis—70.6%
Surface Evolver Bench26.9%—
Epoch Capabilities Index—140.34
ForecastBench—59.6
LiveBench—75.9%

Math Too close to call

Mistral Medium: 28.1 (#245), o3-mini: 28.1 (#244)

Math benchmarks
BenchmarkMistral Mediumo3-mini
OTIS Mock AIME 2024-202532.2%76.9%
LMArena Math14081396
MATH Level 581.6%96.5%
FrontierMath (Feb 2025 set)0.3%12.4%
FrontierMath (Tiers 1-3)—18.6%
FrontierMath Tier 4—0%
ProofBench9%—
LiveBench Math—77.3%
FrontierMath Tier 4 (v1)—4.2%

Knowledge o3-mini leads

Mistral Medium: 25.0 (#265), o3-mini: 38.3 (#146)

Knowledge benchmarks
BenchmarkMistral Mediumo3-mini
GPQA Diamond59.5%77%
LMArena Expert14081364
Humanity's Last Exam4.5%—
SimpleQA Verified—15.3%
Confabulations—17.9%
Vectara Hallucination Rate22.7%—

Multimodal Not comparable

Mistral Medium: 35.3 (#88), o3-mini: —

Multimodal benchmarks
BenchmarkMistral Mediumo3-mini
LMArena Vision1172—

Multilingual Mistral Medium leads

Mistral Medium: 52.1 (#91), o3-mini: 45.7 (#164)

Multilingual benchmarks
BenchmarkMistral Mediumo3-mini
LMArena Non-English14081319
LMArena Chinese14471379
LMArena French14591334
LMArena German14321303
LMArena Japanese13781286
LMArena Korean13801314
LMArena Russian14111304
LMArena Spanish14331321

Instruction Following o3-mini leads

Mistral Medium: 73.7 (#116), o3-mini: 75.1 (#72)

Instruction Following benchmarks
BenchmarkMistral Mediumo3-mini
LMArena Instruction Following13981337
LiveBench Instruction Following—84.4%

Long Context Mistral Medium leads

Mistral Medium: 42.9 (#114), o3-mini: 33.8 (#256)

Long Context benchmarks
BenchmarkMistral Mediumo3-mini
LMArena Longer Query14061343
Fiction.LiveBench—50%

Writing & Preference Mistral Medium leads

Mistral Medium: 60.0 (#103), o3-mini: 50.3 (#182)

Writing & Preference benchmarks
BenchmarkMistral Mediumo3-mini
LMArena Text14241337
LMArena Creative Writing13911286
Short-Story Creative Writing77.3%61.7%
LMArena Multi-Turn14181320
LiveBench Language—50.7%

Frequently asked questions

Is Mistral Medium better than o3-mini?

Mistral Medium and o3-mini score almost the same on the Noometry Index (36.3 vs 36.7), so choose on price, context window or the category you care about most.

Which is cheaper, Mistral Medium or o3-mini?

o3-mini is cheaper. It lists at $1.10 per million input tokens and $4.40 per million output tokens; Mistral Medium lists at $1.50 and $7.50.

Is Mistral Medium or o3-mini better for coding?

o3-mini scores higher on coding benchmarks: 40.8 versus 34.2 in the Noometry coding category.

Which has the bigger context window?

Mistral Medium does, with 262K tokens against 200K.

How many benchmarks do Mistral Medium and o3-mini share?

27 benchmarks have published results for both models. Mistral Medium has 36 scored results on Noometry and o3-mini has 51.

Related comparisons

Go deeper