Model comparison

Mistral Medium vs Qwen Max

Mistral Medium is the stronger model overall, scoring 36.3 to 34.7 on the Noometry Index.

Last verified . 21 shared benchmarks.

Mistral Medium Mistral AI

36.3

Rank #218 Confirmed

Qwen Max Alibaba (Qwen)

34.7

Rank #230 Confirmed

Summary

  • They share 21 benchmarks with published results for both. Mistral Medium scores higher in 6 categories and Qwen Max in 2 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Mistral Medium leads 60.0 to 47.8.
  • The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 32.2% for Mistral Medium and 16.1% for Qwen Max.
  • Qwen Max is cheaper at $1.60 / $6.40 per million input/output tokens, against $1.50 / $7.50 for Mistral Medium.
  • Mistral Medium accepts more context: 262K tokens versus 33K.
  • Mistral Medium has downloadable open weights; the other is API-only.

Side by side

Mistral Medium and Qwen Max specifications
Mistral MediumQwen Max
ProviderMistral AIAlibaba (Qwen)
Noometry Index36.334.7
Released2023-12-112024-04-03
WeightsOpenProprietary
Context window262K33K
Max output262K8K
Input $ / M tokens$1.50$1.60
Output $ / M tokens$7.50$6.40
Results tracked3623

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Mistral Medium leads

Mistral Medium: 34.2 (#243), Qwen Max: 30.7 (#292)

Coding benchmarks
BenchmarkMistral MediumQwen Max
LMArena Coding14341288
FrontierCode8%—
Aider Polyglot—21.8%
SciCode40.2%—
WeirdML43.7%—
ALE-Bench763.98—

Agentic & Tool Use Not comparable

Mistral Medium: 28.3 (#90), Qwen Max: —

Agentic & Tool Use benchmarks
BenchmarkMistral MediumQwen Max
Berkeley Function Calling Leaderboard37.7%—

Reasoning Qwen Max leads

Mistral Medium: 24.0 (#167), Qwen Max: 25.1 (#151)

Reasoning benchmarks
BenchmarkMistral MediumQwen Max
LMArena Hard Prompts14261269
Kagi LLM Benchmark50%—
CritPt0%—
DTBench75.5%—
LMCA26.1%—
Surface Evolver Bench26.9%—

Math Mistral Medium leads

Mistral Medium: 28.1 (#245), Qwen Max: 22.3 (#276)

Math benchmarks
BenchmarkMistral MediumQwen Max
OTIS Mock AIME 2024-202532.2%16.1%
LMArena Math14081275
MATH Level 581.6%67.2%
FrontierMath (Feb 2025 set)0.3%1%
ProofBench9%—

Knowledge Qwen Max leads

Mistral Medium: 25.0 (#265), Qwen Max: 30.3 (#228)

Knowledge benchmarks
BenchmarkMistral MediumQwen Max
GPQA Diamond59.5%56.1%
LMArena Expert14081248
Humanity's Last Exam4.5%—
Vectara Hallucination Rate22.7%—

Multimodal Not comparable

Mistral Medium: 35.3 (#88), Qwen Max: —

Multimodal benchmarks
BenchmarkMistral MediumQwen Max
LMArena Vision1172—

Multilingual Mistral Medium leads

Mistral Medium: 52.1 (#91), Qwen Max: 41.8 (#202)

Multilingual benchmarks
BenchmarkMistral MediumQwen Max
LMArena Non-English14081263
LMArena Chinese14471254
LMArena French14591330
LMArena German14321254
LMArena Japanese13781205
LMArena Korean13801142
LMArena Russian14111274
LMArena Spanish14331290

Instruction Following Mistral Medium leads

Mistral Medium: 73.7 (#116), Qwen Max: 66.5 (#208)

Instruction Following benchmarks
BenchmarkMistral MediumQwen Max
LMArena Instruction Following13981262

Long Context Mistral Medium leads

Mistral Medium: 42.9 (#114), Qwen Max: 39.4 (#180)

Long Context benchmarks
BenchmarkMistral MediumQwen Max
LMArena Longer Query14061288
Fiction.LiveBench—66.7%

Writing & Preference Mistral Medium leads

Mistral Medium: 60.0 (#103), Qwen Max: 47.8 (#205)

Writing & Preference benchmarks
BenchmarkMistral MediumQwen Max
LMArena Text14241282
LMArena Creative Writing13911248
LMArena Multi-Turn14181277
Short-Story Creative Writing77.3%—

Frequently asked questions

Is Mistral Medium better than Qwen Max?

Mistral Medium is the stronger model overall, scoring 36.3 to 34.7 on the Noometry Index.

Which is cheaper, Mistral Medium or Qwen Max?

Qwen Max is cheaper. It lists at $1.60 per million input tokens and $6.40 per million output tokens; Mistral Medium lists at $1.50 and $7.50.

Is Mistral Medium or Qwen Max better for coding?

Mistral Medium scores higher on coding benchmarks: 34.2 versus 30.7 in the Noometry coding category.

Which has the bigger context window?

Mistral Medium does, with 262K tokens against 33K.

How many benchmarks do Mistral Medium and Qwen Max share?

21 benchmarks have published results for both models. Mistral Medium has 36 scored results on Noometry and Qwen Max has 23.

Related comparisons

Go deeper