Model comparison

Mercury 2.5 vs Mistral Small

Mercury 2.5 and Mistral Small score almost the same on the Noometry Index (33.5 vs 33.4), so choose on price, context window or the category you care about most.

Last verified . 3 shared benchmarks.

Mercury 2.5 Inception

33.5

Rank #242 Reported

Mistral Small Mistral AI

33.4

Rank #243 Confirmed

Summary

  • They share 3 benchmarks with published results for both. Mercury 2.5 scores higher in 3 categories and Mistral Small in 0 categories; 3 gaps are clear of the uncertainty.
  • The widest gap is in math, where Mercury 2.5 leads 23.3 to 16.4.
  • The biggest single-benchmark swing is SciCode: 38.5% for Mercury 2.5 and 26.5% for Mistral Small.
  • Mercury 2.5 is cheaper at $0.04 / $0.15 per million input/output tokens, against $0.15 / $0.60 for Mistral Small.
  • Mistral Small accepts more context: 262K tokens versus 260K.
  • Mistral Small has downloadable open weights; the other is API-only.

Side by side

Mercury 2.5 and Mistral Small specifications
Mercury 2.5Mistral Small
ProviderInceptionMistral AI
Noometry Index33.533.4
Released2026-09-082024-02-26
WeightsProprietaryOpen
Context window260K262K
Max output66K256K
Input $ / M tokens$0.04$0.15
Output $ / M tokens$0.15$0.60
Results tracked439

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Mercury 2.5 leads

Mercury 2.5: 39.5 (#156), Mistral Small: 34.0 (#247)

Coding benchmarks
BenchmarkMercury 2.5Mistral Small
SciCode38.5%26.5%
ALE-Bench301.65497.62
BigCodeBench Instruct—36.1%
LiveBench Coding—36.2%
LMArena Coding—1362
BigCodeBench Complete—46.6%

Agentic & Tool Use Not comparable

Mercury 2.5: —, Mistral Small: 28.1 (#93)

Agentic & Tool Use benchmarks
BenchmarkMercury 2.5Mistral Small
Berkeley Function Calling Leaderboard—37.1%

Reasoning Mercury 2.5 leads

Mercury 2.5: 22.4 (#193), Mistral Small: 19.8 (#250)

Reasoning benchmarks
BenchmarkMercury 2.5Mistral Small
CritPt0%0%
Kagi LLM Benchmark—37.8%
LiveBench Reasoning—44.8%
LMArena Hard Prompts—1335
DTBench—70.9%
LiveBench Data Analysis—53.7%
LMCA—20.6%
LiveBench—44%

Math Mercury 2.5 leads

Mercury 2.5: 23.3 (#272), Mistral Small: 16.4 (#293)

Math benchmarks
BenchmarkMercury 2.5Mistral Small
OTIS Mock AIME 2024-2025—5.8%
ProofBench3%—
LiveBench Math—39.9%
LMArena Math—1341
MATH Level 5—46.8%

Knowledge Not comparable

Mercury 2.5: —, Mistral Small: 31.0 (#222)

Knowledge benchmarks
BenchmarkMercury 2.5Mistral Small
GPQA Diamond—47.5%
Vectara Hallucination Rate—5.1%
LMArena Expert—1291
MMLU—68.7%

Multimodal Not comparable

Mercury 2.5: —, Mistral Small: 33.5 (#96)

Multimodal benchmarks
BenchmarkMercury 2.5Mistral Small
LMArena Vision—1142

Multilingual Not comparable

Mercury 2.5: —, Mistral Small: 45.5 (#169)

Multilingual benchmarks
BenchmarkMercury 2.5Mistral Small
LMArena Non-English—1315
LMArena Chinese—1340
LMArena French—1337
LMArena German—1340
LMArena Japanese—1275
LMArena Korean—1259
LMArena Russian—1324
LMArena Spanish—1346

Instruction Following Not comparable

Mercury 2.5: —, Mistral Small: 66.4 (#209)

Instruction Following benchmarks
BenchmarkMercury 2.5Mistral Small
LiveBench Instruction Following—63.7%
LMArena Instruction Following—1310

Long Context Not comparable

Mercury 2.5: —, Mistral Small: 40.4 (#156)

Long Context benchmarks
BenchmarkMercury 2.5Mistral Small
LMArena Longer Query—1327

Writing & Preference Not comparable

Mercury 2.5: —, Mistral Small: 52.5 (#171)

Writing & Preference benchmarks
BenchmarkMercury 2.5Mistral Small
LMArena Text—1338
LMArena Creative Writing—1305
LMArena Multi-Turn—1344
LiveBench Language—30.5%

Frequently asked questions

Is Mercury 2.5 better than Mistral Small?

Mercury 2.5 and Mistral Small score almost the same on the Noometry Index (33.5 vs 33.4), so choose on price, context window or the category you care about most.

Which is cheaper, Mercury 2.5 or Mistral Small?

Mercury 2.5 is cheaper. It lists at $0.04 per million input tokens and $0.15 per million output tokens; Mistral Small lists at $0.15 and $0.60.

Is Mercury 2.5 or Mistral Small better for coding?

Mercury 2.5 scores higher on coding benchmarks: 39.5 versus 34.0 in the Noometry coding category.

Which has the bigger context window?

Mistral Small does, with 262K tokens against 260K.

How many benchmarks do Mercury 2.5 and Mistral Small share?

3 benchmarks have published results for both models. Mercury 2.5 has 4 scored results on Noometry and Mistral Small has 39.

Related comparisons

Go deeper