Model comparison

Mercury 2.5 vs Mistral Large 3

Mistral Large 3 is the stronger model overall, scoring 39.1 to 33.5 on the Noometry Index. Mercury 2.5 costs 5.6× less per token, which makes it the better buy when Mistral Large 3's lead doesn't matter for your workload.

Last verified . 0 shared benchmarks.

Mercury 2.5 Inception

33.5

Rank #242 Reported

Mistral Large 3 Mistral AI

39.1

Rank #176 Confirmed

Summary

  • The widest gap is in math, where Mistral Large 3 leads 38.7 to 23.3.
  • Mercury 2.5 is cheaper at $0.04 / $0.15 per million input/output tokens, against $0.25 / $0.75 for Mistral Large 3.
  • Mistral Large 3 accepts more context: 262K tokens versus 260K.
  • Mistral Large 3 has downloadable open weights; the other is API-only.

Side by side

Mercury 2.5 and Mistral Large 3 specifications
Mercury 2.5Mistral Large 3
ProviderInceptionMistral AI
Noometry Index33.539.1
Released2026-09-082025-12-02
WeightsProprietaryOpen
Context window260K262K
Max output66K8K
Input $ / M tokens$0.04$0.25
Output $ / M tokens$0.15$0.75
Results tracked424

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Mercury 2.5 leads

Mercury 2.5: 39.5 (#156), Mistral Large 3: 34.4 (#237)

Coding benchmarks
BenchmarkMercury 2.5Mistral Large 3
LMArena WebDev—1230
SciCode38.5%—
LMArena Coding—1448
ALE-Bench301.65—

Reasoning Mercury 2.5 leads

Mercury 2.5: 22.4 (#193), Mistral Large 3: 15.2 (#319)

Reasoning benchmarks
BenchmarkMercury 2.5Mistral Large 3
Kagi LLM Benchmark—50.9%
NYT Connections (extended)—7.5%
CritPt0%—
Thematic Generalization—23%
LMArena Hard Prompts—1429

Math Mistral Large 3 leads

Mercury 2.5: 23.3 (#272), Mistral Large 3: 38.7 (#129)

Math benchmarks
BenchmarkMercury 2.5Mistral Large 3
ProofBench3%—
LMArena Math—1414

Knowledge Not comparable

Mercury 2.5: —, Mistral Large 3: 36.0 (#177)

Knowledge benchmarks
BenchmarkMercury 2.5Mistral Large 3
Vectara Hallucination Rate—14.5%
LMArena Expert—1421

Multimodal Not comparable

Mercury 2.5: —, Mistral Large 3: 38.2 (#66)

Multimodal benchmarks
BenchmarkMercury 2.5Mistral Large 3
LMArena Vision—1221

Multilingual Not comparable

Mercury 2.5: —, Mistral Large 3: 52.5 (#84)

Multilingual benchmarks
BenchmarkMercury 2.5Mistral Large 3
LMArena Non-English—1413
LMArena Chinese—1447
LMArena French—1455
LMArena German—1437
LMArena Japanese—1394
LMArena Korean—1384
LMArena Russian—1411
LMArena Spanish—1440

Instruction Following Not comparable

Mercury 2.5: —, Mistral Large 3: 74.0 (#108)

Instruction Following benchmarks
BenchmarkMercury 2.5Mistral Large 3
LMArena Instruction Following—1403

Long Context Not comparable

Mercury 2.5: —, Mistral Large 3: 43.1 (#105)

Long Context benchmarks
BenchmarkMercury 2.5Mistral Large 3
LMArena Longer Query—1413

Writing & Preference Not comparable

Mercury 2.5: —, Mistral Large 3: 60.0 (#101)

Writing & Preference benchmarks
BenchmarkMercury 2.5Mistral Large 3
LMArena Text—1428
LMArena Creative Writing—1386
EQ-Bench Creative Writing—1412
LMArena Multi-Turn—1429

Frequently asked questions

Is Mercury 2.5 better than Mistral Large 3?

Mistral Large 3 is the stronger model overall, scoring 39.1 to 33.5 on the Noometry Index. Mercury 2.5 costs 5.6× less per token, which makes it the better buy when Mistral Large 3's lead doesn't matter for your workload.

Which is cheaper, Mercury 2.5 or Mistral Large 3?

Mercury 2.5 is cheaper. It lists at $0.04 per million input tokens and $0.15 per million output tokens; Mistral Large 3 lists at $0.25 and $0.75.

Is Mercury 2.5 or Mistral Large 3 better for coding?

Mercury 2.5 scores higher on coding benchmarks: 39.5 versus 34.4 in the Noometry coding category.

Which has the bigger context window?

Mistral Large 3 does, with 262K tokens against 260K.

How many benchmarks do Mercury 2.5 and Mistral Large 3 share?

0 benchmarks have published results for both models. Mercury 2.5 has 4 scored results on Noometry and Mistral Large 3 has 24.

Related comparisons

Go deeper