Model comparison

Mercury 2.5 vs Qwen2.5-Coder-32B

Mercury 2.5 and Qwen2.5-Coder-32B score almost the same on the Noometry Index (33.5 vs 33.4), so choose on price, context window or the category you care about most.

Last verified . 0 shared benchmarks.

Mercury 2.5 Inception

33.5

Rank #242 Reported

Qwen2.5-Coder-32B Alibaba (Qwen)

33.4

Rank #245 Confirmed

Summary

  • The widest gap is in coding, where Mercury 2.5 leads 39.5 to 22.6.
  • Mercury 2.5 is cheaper at $0.04 / $0.15 per million input/output tokens, against $0.66 / $1 for Qwen2.5-Coder-32B.
  • Mercury 2.5 accepts more context: 260K tokens versus 33K.
  • Qwen2.5-Coder-32B has downloadable open weights; the other is API-only.

Side by side

Mercury 2.5 and Qwen2.5-Coder-32B specifications
Mercury 2.5Qwen2.5-Coder-32B
ProviderInceptionAlibaba (Qwen)
Noometry Index33.533.4
Released2026-09-082024-09-18
WeightsProprietaryOpen
Context window260K33K
Max output66K29K
Input $ / M tokens$0.04$0.66
Output $ / M tokens$0.15$1
Results tracked431

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Mercury 2.5 leads

Mercury 2.5: 39.5 (#156), Qwen2.5-Coder-32B: 22.6 (#333)

Coding benchmarks
BenchmarkMercury 2.5Qwen2.5-Coder-32B
SWE-bench Verified (bash only)—9%
Aider Polyglot—16.4%
SciCode38.5%—
BigCodeBench Instruct—49%
LiveBench Coding—56.9%
LMArena Coding—1276
BigCodeBench Complete—58%
ALE-Bench301.65—
HumanEval+—87.2%
MBPP+—77%

Reasoning Mercury 2.5 leads

Mercury 2.5: 22.4 (#193), Qwen2.5-Coder-32B: 21.2 (#225)

Reasoning benchmarks
BenchmarkMercury 2.5Qwen2.5-Coder-32B
CritPt0%—
LiveBench Reasoning—42.1%
LMArena Hard Prompts—1251
LiveBench Data Analysis—49.9%
Epoch Capabilities Index—119.49
HellaSwag—83%
LiveBench—46.2%
WinoGrande—80.8%

Math Qwen2.5-Coder-32B leads

Mercury 2.5: 23.3 (#272), Qwen2.5-Coder-32B: 33.3 (#204)

Math benchmarks
BenchmarkMercury 2.5Qwen2.5-Coder-32B
ProofBench3%—
LiveBench Math—46.6%
LMArena Math—1251
GSM8K—93%

Knowledge Not comparable

Mercury 2.5: —, Qwen2.5-Coder-32B: 33.4 (#203)

Knowledge benchmarks
BenchmarkMercury 2.5Qwen2.5-Coder-32B
LMArena Expert—1221
ARC (AI2) Challenge—70.5%
MMLU—79.1%

Multilingual Not comparable

Mercury 2.5: —, Qwen2.5-Coder-32B: 37.8 (#235)

Multilingual benchmarks
BenchmarkMercury 2.5Qwen2.5-Coder-32B
LMArena Non-English—1205
LMArena Chinese—1222
LMArena Russian—1228

Instruction Following Not comparable

Mercury 2.5: —, Qwen2.5-Coder-32B: 61.4 (#245)

Instruction Following benchmarks
BenchmarkMercury 2.5Qwen2.5-Coder-32B
LiveBench Instruction Following—58.7%
LMArena Instruction Following—1223

Long Context Not comparable

Mercury 2.5: —, Qwen2.5-Coder-32B: 38.0 (#208)

Long Context benchmarks
BenchmarkMercury 2.5Qwen2.5-Coder-32B
LMArena Longer Query—1251

Writing & Preference Not comparable

Mercury 2.5: —, Qwen2.5-Coder-32B: 41.6 (#240)

Writing & Preference benchmarks
BenchmarkMercury 2.5Qwen2.5-Coder-32B
LMArena Text—1230
LMArena Creative Writing—1174
LMArena Multi-Turn—1222
LiveBench Language—23.3%

Frequently asked questions

Is Mercury 2.5 better than Qwen2.5-Coder-32B?

Mercury 2.5 and Qwen2.5-Coder-32B score almost the same on the Noometry Index (33.5 vs 33.4), so choose on price, context window or the category you care about most.

Which is cheaper, Mercury 2.5 or Qwen2.5-Coder-32B?

Mercury 2.5 is cheaper. It lists at $0.04 per million input tokens and $0.15 per million output tokens; Qwen2.5-Coder-32B lists at $0.66 and $1.

Is Mercury 2.5 or Qwen2.5-Coder-32B better for coding?

Mercury 2.5 scores higher on coding benchmarks: 39.5 versus 22.6 in the Noometry coding category.

Which has the bigger context window?

Mercury 2.5 does, with 260K tokens against 33K.

How many benchmarks do Mercury 2.5 and Qwen2.5-Coder-32B share?

0 benchmarks have published results for both models. Mercury 2.5 has 4 scored results on Noometry and Qwen2.5-Coder-32B has 31.

Related comparisons

Go deeper