Model comparison

Mercury vs MiniMax-M2

Mercury and MiniMax-M2 score almost the same on the Noometry Index (37.6 vs 37.4), so choose on price, context window or the category you care about most.

Last verified . 9 shared benchmarks.

Mercury Inception

37.6

Rank #199 Confirmed

MiniMax-M2 MiniMax

37.4

Rank #204 Confirmed

Summary

  • They share 9 benchmarks with published results for both. Mercury scores higher in 0 categories and MiniMax-M2 in 6 categories; 5 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where MiniMax-M2 leads 53.0 to 46.2.
  • The biggest single-benchmark swing is Kagi LLM Benchmark: 21.6% for Mercury and 57.8% for MiniMax-M2.
  • MiniMax-M2 has downloadable open weights; the other is API-only.

Side by side

Mercury and MiniMax-M2 specifications
MercuryMiniMax-M2
ProviderInceptionMiniMax
Noometry Index37.637.4
Released—2025-10-27
WeightsProprietaryOpen
Context window—205K
Max output—131K
Input $ / M tokens—$0.30
Output $ / M tokens—$1.20
Results tracked921

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

Mercury: 38.7 (#170), MiniMax-M2: 39.3 (#159)

Coding benchmarks
BenchmarkMercuryMiniMax-M2
LMArena Coding13221370
SWE-bench Verified (bash only)—61%
LMArena WebDev—1297

Agentic & Tool Use Not comparable

Mercury: —, MiniMax-M2: 25.1 (#109)

Agentic & Tool Use benchmarks
BenchmarkMercuryMiniMax-M2
Terminal-Bench—30%
Vending-Bench 2—160.6

Reasoning MiniMax-M2 leads

Mercury: 17.5 (#293), MiniMax-M2: 19.4 (#258)

Reasoning benchmarks
BenchmarkMercuryMiniMax-M2
Kagi LLM Benchmark21.6%57.8%
LMArena Hard Prompts12851357
NYT Connections (extended)—14.8%

Math Not comparable

Mercury: —, MiniMax-M2: 37.3 (#160)

Math benchmarks
BenchmarkMercuryMiniMax-M2
LMArena Math—1352

Knowledge Not comparable

Mercury: —, MiniMax-M2: 37.0 (#163)

Knowledge benchmarks
BenchmarkMercuryMiniMax-M2
LMArena Expert—1337

Multilingual MiniMax-M2 leads

Mercury: 41.6 (#206), MiniMax-M2: 45.3 (#171)

Multilingual benchmarks
BenchmarkMercuryMiniMax-M2
LMArena Non-English12601313
LMArena Chinese—1366
LMArena French—1335
LMArena German—1355
LMArena Russian—1331
LMArena Spanish—1326

Instruction Following MiniMax-M2 leads

Mercury: 65.2 (#224), MiniMax-M2: 70.2 (#166)

Instruction Following benchmarks
BenchmarkMercuryMiniMax-M2
LMArena Instruction Following12391328

Long Context MiniMax-M2 leads

Mercury: 38.4 (#198), MiniMax-M2: 40.5 (#153)

Long Context benchmarks
BenchmarkMercuryMiniMax-M2
LMArena Longer Query12661331

Writing & Preference MiniMax-M2 leads

Mercury: 46.2 (#221), MiniMax-M2: 53.0 (#162)

Writing & Preference benchmarks
BenchmarkMercuryMiniMax-M2
LMArena Text12821340
LMArena Creative Writing11911286
LMArena Multi-Turn12821361

Frequently asked questions

Is Mercury better than MiniMax-M2?

Mercury and MiniMax-M2 score almost the same on the Noometry Index (37.6 vs 37.4), so choose on price, context window or the category you care about most.

Is Mercury or MiniMax-M2 better for coding?

They score almost the same on coding (38.7 vs 39.3); test both on your own repository before choosing.

How many benchmarks do Mercury and MiniMax-M2 share?

9 benchmarks have published results for both models. Mercury has 9 scored results on Noometry and MiniMax-M2 has 21.

Related comparisons

Go deeper