Model comparison

Qwen2.5-Coder-32B vs Wizardlm 70b

Qwen2.5-Coder-32B and Wizardlm 70b score almost the same on the Noometry Index (33.4 vs 33.0), so choose on price, context window or the category you care about most.

Last verified . 11 shared benchmarks.

Qwen2.5-Coder-32B Alibaba (Qwen)

33.4

Rank #245 Confirmed

Wizardlm 70b Microsoft

33.0

Rank #249 Confirmed

Summary

  • They share 11 benchmarks with published results for both. Qwen2.5-Coder-32B scores higher in 6 categories and Wizardlm 70b in 1 category; 6 gaps are clear of the uncertainty.
  • The widest gap is in coding, where Wizardlm 70b leads 31.4 to 22.6.

Side by side

Qwen2.5-Coder-32B and Wizardlm 70b specifications
Qwen2.5-Coder-32BWizardlm 70b
ProviderAlibaba (Qwen)Microsoft
Noometry Index33.433.0
Released2024-09-18—
WeightsOpenOpen
Context window33K—
Max output29K—
Input $ / M tokens$0.66—
Output $ / M tokens$1—
Results tracked3112

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Wizardlm 70b leads

Qwen2.5-Coder-32B: 22.6 (#333), Wizardlm 70b: 31.4 (#285)

Coding benchmarks
BenchmarkQwen2.5-Coder-32BWizardlm 70b
LMArena Coding12761081
SWE-bench Verified (bash only)9%—
Aider Polyglot16.4%—
BigCodeBench Instruct49%—
LiveBench Coding56.9%—
BigCodeBench Complete58%—
HumanEval+87.2%—
MBPP+77%—

Reasoning Too close to call

Qwen2.5-Coder-32B: 21.2 (#225), Wizardlm 70b: 20.7 (#234)

Reasoning benchmarks
BenchmarkQwen2.5-Coder-32BWizardlm 70b
LMArena Hard Prompts12511079
LiveBench Reasoning42.1%—
LiveBench Data Analysis49.9%—
Epoch Capabilities Index119.49—
HellaSwag83%—
LiveBench46.2%—
WinoGrande80.8%—

Math Qwen2.5-Coder-32B leads

Qwen2.5-Coder-32B: 33.3 (#204), Wizardlm 70b: 32.2 (#218)

Math benchmarks
BenchmarkQwen2.5-Coder-32BWizardlm 70b
LMArena Math12511116
LiveBench Math46.6%—
GSM8K93%—

Knowledge Not comparable

Qwen2.5-Coder-32B: 33.4 (#203), Wizardlm 70b: —

Knowledge benchmarks
BenchmarkQwen2.5-Coder-32BWizardlm 70b
LMArena Expert1221—
ARC (AI2) Challenge70.5%—
MMLU79.1%—

Multilingual Qwen2.5-Coder-32B leads

Qwen2.5-Coder-32B: 37.8 (#235), Wizardlm 70b: 29.6 (#265)

Multilingual benchmarks
BenchmarkQwen2.5-Coder-32BWizardlm 70b
LMArena Non-English12051078
LMArena Chinese12221052
LMArena Russian12281155
LMArena German—1083

Instruction Following Qwen2.5-Coder-32B leads

Qwen2.5-Coder-32B: 61.4 (#245), Wizardlm 70b: 56.3 (#273)

Instruction Following benchmarks
BenchmarkQwen2.5-Coder-32BWizardlm 70b
LMArena Instruction Following12231093
LiveBench Instruction Following58.7%—

Long Context Qwen2.5-Coder-32B leads

Qwen2.5-Coder-32B: 38.0 (#208), Wizardlm 70b: 33.3 (#263)

Long Context benchmarks
BenchmarkQwen2.5-Coder-32BWizardlm 70b
LMArena Longer Query12511097

Writing & Preference Qwen2.5-Coder-32B leads

Qwen2.5-Coder-32B: 41.6 (#240), Wizardlm 70b: 34.8 (#269)

Writing & Preference benchmarks
BenchmarkQwen2.5-Coder-32BWizardlm 70b
LMArena Text12301120
LMArena Creative Writing11741149
LMArena Multi-Turn12221108
LiveBench Language23.3%—

Frequently asked questions

Is Qwen2.5-Coder-32B better than Wizardlm 70b?

Qwen2.5-Coder-32B and Wizardlm 70b score almost the same on the Noometry Index (33.4 vs 33.0), so choose on price, context window or the category you care about most.

Is Qwen2.5-Coder-32B or Wizardlm 70b better for coding?

Wizardlm 70b scores higher on coding benchmarks: 31.4 versus 22.6 in the Noometry coding category.

How many benchmarks do Qwen2.5-Coder-32B and Wizardlm 70b share?

11 benchmarks have published results for both models. Qwen2.5-Coder-32B has 31 scored results on Noometry and Wizardlm 70b has 12.

Related comparisons

Go deeper