Model comparison

Qwen2.5-Coder-32B vs Wizardlm 13b

Qwen2.5-Coder-32B is the stronger model overall, scoring 33.4 to 31.4 on the Noometry Index.

Last verified . 10 shared benchmarks.

Qwen2.5-Coder-32B Alibaba (Qwen)

33.4

Rank #245 Confirmed

Wizardlm 13b Microsoft

31.4

Rank #274 Confirmed

Summary

  • They share 10 benchmarks with published results for both. Qwen2.5-Coder-32B scores higher in 6 categories and Wizardlm 13b in 1 category; 7 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Qwen2.5-Coder-32B leads 41.6 to 30.5.

Side by side

Qwen2.5-Coder-32B and Wizardlm 13b specifications
Qwen2.5-Coder-32BWizardlm 13b
ProviderAlibaba (Qwen)Microsoft
Noometry Index33.431.4
Released2024-09-18—
WeightsOpenOpen
Context window33K—
Max output29K—
Input $ / M tokens$0.66—
Output $ / M tokens$1—
Results tracked3110

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Wizardlm 13b leads

Qwen2.5-Coder-32B: 22.6 (#333), Wizardlm 13b: 30.1 (#298)

Coding benchmarks
BenchmarkQwen2.5-Coder-32BWizardlm 13b
LMArena Coding12761035
SWE-bench Verified (bash only)9%—
Aider Polyglot16.4%—
BigCodeBench Instruct49%—
LiveBench Coding56.9%—
BigCodeBench Complete58%—
HumanEval+87.2%—
MBPP+77%—

Reasoning Qwen2.5-Coder-32B leads

Qwen2.5-Coder-32B: 21.2 (#225), Wizardlm 13b: 19.4 (#259)

Reasoning benchmarks
BenchmarkQwen2.5-Coder-32BWizardlm 13b
LMArena Hard Prompts12511018
LiveBench Reasoning42.1%—
LiveBench Data Analysis49.9%—
Epoch Capabilities Index119.49—
HellaSwag83%—
LiveBench46.2%—
WinoGrande80.8%—

Math Qwen2.5-Coder-32B leads

Qwen2.5-Coder-32B: 33.3 (#204), Wizardlm 13b: 30.2 (#238)

Math benchmarks
BenchmarkQwen2.5-Coder-32BWizardlm 13b
LMArena Math12511017
LiveBench Math46.6%—
GSM8K93%—

Knowledge Not comparable

Qwen2.5-Coder-32B: 33.4 (#203), Wizardlm 13b: —

Knowledge benchmarks
BenchmarkQwen2.5-Coder-32BWizardlm 13b
LMArena Expert1221—
ARC (AI2) Challenge70.5%—
MMLU79.1%—

Multilingual Qwen2.5-Coder-32B leads

Qwen2.5-Coder-32B: 37.8 (#235), Wizardlm 13b: 27.1 (#277)

Multilingual benchmarks
BenchmarkQwen2.5-Coder-32BWizardlm 13b
LMArena Non-English12051034
LMArena Chinese12221023
LMArena Russian1228—

Instruction Following Qwen2.5-Coder-32B leads

Qwen2.5-Coder-32B: 61.4 (#245), Wizardlm 13b: 53.5 (#285)

Instruction Following benchmarks
BenchmarkQwen2.5-Coder-32BWizardlm 13b
LMArena Instruction Following12231048
LiveBench Instruction Following58.7%—

Long Context Qwen2.5-Coder-32B leads

Qwen2.5-Coder-32B: 38.0 (#208), Wizardlm 13b: 32.0 (#273)

Long Context benchmarks
BenchmarkQwen2.5-Coder-32BWizardlm 13b
LMArena Longer Query12511054

Writing & Preference Qwen2.5-Coder-32B leads

Qwen2.5-Coder-32B: 41.6 (#240), Wizardlm 13b: 30.5 (#287)

Writing & Preference benchmarks
BenchmarkQwen2.5-Coder-32BWizardlm 13b
LMArena Text12301077
LMArena Creative Writing11741091
LMArena Multi-Turn12221047
LiveBench Language23.3%—

Frequently asked questions

Is Qwen2.5-Coder-32B better than Wizardlm 13b?

Qwen2.5-Coder-32B is the stronger model overall, scoring 33.4 to 31.4 on the Noometry Index.

Is Qwen2.5-Coder-32B or Wizardlm 13b better for coding?

Wizardlm 13b scores higher on coding benchmarks: 30.1 versus 22.6 in the Noometry coding category.

How many benchmarks do Qwen2.5-Coder-32B and Wizardlm 13b share?

10 benchmarks have published results for both models. Qwen2.5-Coder-32B has 31 scored results on Noometry and Wizardlm 13b has 10.

Related comparisons

Go deeper