Model comparison

Mistral Large vs Qwen2.5 72B Instruct

Mistral Large and Qwen2.5 72B Instruct score almost the same on the Noometry Index (31.9 vs 31.9), so choose on price, context window or the category you care about most.

Last verified . 33 shared benchmarks.

Mistral Large Mistral AI

31.9

Rank #263 Confirmed

Qwen2.5 72B Instruct Alibaba (Qwen)

31.9

Rank #267 Confirmed

Summary

  • They share 33 benchmarks with published results for both. Mistral Large scores higher in 4 categories and Qwen2.5 72B Instruct in 5 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Qwen2.5 72B Instruct leads 22.3 to 15.8.
  • The biggest single-benchmark swing is BigCodeBench Complete: 38.3% for Mistral Large and 55.9% for Qwen2.5 72B Instruct.
  • Qwen2.5 72B Instruct is cheaper at $1.40 / $5.60 per million input/output tokens, against $2 / $6 for Mistral Large.

Side by side

Mistral Large and Qwen2.5 72B Instruct specifications
Mistral LargeQwen2.5 72B Instruct
ProviderMistral AIAlibaba (Qwen)
Noometry Index31.931.9
Released2024-02-262024-09
WeightsOpenOpen
Context window131K131K
Max output16K8K
Input $ / M tokens$2$1.40
Output $ / M tokens$6$5.60
Results tracked5143

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Mistral Large leads

Mistral Large: 34.3 (#240), Qwen2.5 72B Instruct: 33.2 (#260)

Coding benchmarks
BenchmarkMistral LargeQwen2.5 72B Instruct
BigCodeBench Instruct30%45.8%
LMArena Coding12771292
BigCodeBench Complete38.3%55.9%
SciCode36.2%—
WeirdML—16%
LiveBench Coding47.1%—
ALE-Bench264.7—
HumanEval+62.2%—
MBPP+59.5%—

Agentic & Tool Use Mistral Large leads

Mistral Large: 28.6 (#89), Qwen2.5 72B Instruct: 22.1 (#133)

Agentic & Tool Use benchmarks
BenchmarkMistral LargeQwen2.5 72B Instruct
Berkeley Function Calling Leaderboard38.4%—
TheAgentCompany—5.7%
BALROG—16.2%
METR Time Horizons—35.8%

Reasoning Qwen2.5 72B Instruct leads

Mistral Large: 15.8 (#310), Qwen2.5 72B Instruct: 22.3 (#199)

Reasoning benchmarks
BenchmarkMistral LargeQwen2.5 72B Instruct
LMArena Hard Prompts12571271
DTBench65.1%62.9%
LMCA16.7%13.4%
Epoch Capabilities Index128.52129
ForecastBench57.157.5
SimpleBench22.5%—
CritPt0%—
LiveBench Reasoning43.5%—
LiveBench Data Analysis50.1%—
BIG-Bench Hard—79.8%
HellaSwag—84.8%
LiveBench48.4%—
PIQA—82.6%
WinoGrande—82.3%

Math Qwen2.5 72B Instruct leads

Mistral Large: 18.2 (#291), Qwen2.5 72B Instruct: 19.3 (#287)

Math benchmarks
BenchmarkMistral LargeQwen2.5 72B Instruct
OTIS Mock AIME 2024-20258.5%8.1%
Omni-MATH28.1%33%
LMArena Math12621283
MATH Level 550.3%63.2%
LiveBench Math42.5%—
FrontierMath (Feb 2025 set)0.3%—

Knowledge Mistral Large leads

Mistral Large: 30.1 (#230), Qwen2.5 72B Instruct: 27.0 (#253)

Knowledge benchmarks
BenchmarkMistral LargeQwen2.5 72B Instruct
GPQA Diamond51.3%49.1%
MMLU-Pro59.9%63.1%
Confabulations21.4%19.1%
GPQA (HELM)43.5%42.6%
LMArena Expert12321245
MMLU80%85.3%
Vectara Hallucination Rate4.5%—
ARC (AI2) Challenge—94.5%
TriviaQA—71.9%

Multilingual Qwen2.5 72B Instruct leads

Mistral Large: 40.0 (#219), Qwen2.5 72B Instruct: 41.0 (#213)

Multilingual benchmarks
BenchmarkMistral LargeQwen2.5 72B Instruct
LMArena Non-English12371252
LMArena Chinese12401272
LMArena French13251280
LMArena German12541234
LMArena Japanese11881180
LMArena Korean12021188
LMArena Russian12571264
LMArena Spanish12681256

Instruction Following Mistral Large leads

Mistral Large: 67.9 (#191), Qwen2.5 72B Instruct: 65.5 (#221)

Instruction Following benchmarks
BenchmarkMistral LargeQwen2.5 72B Instruct
IFEval87.7%80.6%
LMArena Instruction Following12491254
LiveBench Instruction Following67.9%—

Long Context Too close to call

Mistral Large: 38.3 (#199), Qwen2.5 72B Instruct: 38.9 (#188)

Long Context benchmarks
BenchmarkMistral LargeQwen2.5 72B Instruct
LMArena Longer Query12611282

Writing & Preference Qwen2.5 72B Instruct leads

Mistral Large: 40.7 (#242), Qwen2.5 72B Instruct: 46.7 (#215)

Writing & Preference benchmarks
BenchmarkMistral LargeQwen2.5 72B Instruct
LMArena Text12661269
LMArena Creative Writing12431221
WildBench80.1%80.2%
LMArena Multi-Turn12601272
Short-Story Creative Writing69%—
EQ-Bench Creative Writing985—
LiveBench Language39.4%—

Frequently asked questions

Is Mistral Large better than Qwen2.5 72B Instruct?

Mistral Large and Qwen2.5 72B Instruct score almost the same on the Noometry Index (31.9 vs 31.9), so choose on price, context window or the category you care about most.

Which is cheaper, Mistral Large or Qwen2.5 72B Instruct?

Qwen2.5 72B Instruct is cheaper. It lists at $1.40 per million input tokens and $5.60 per million output tokens; Mistral Large lists at $2 and $6.

Is Mistral Large or Qwen2.5 72B Instruct better for coding?

Mistral Large scores higher on coding benchmarks: 34.3 versus 33.2 in the Noometry coding category.

Which has the bigger context window?

Both accept 131K tokens.

How many benchmarks do Mistral Large and Qwen2.5 72B Instruct share?

33 benchmarks have published results for both models. Mistral Large has 51 scored results on Noometry and Qwen2.5 72B Instruct has 43.

Related comparisons

Go deeper