Model comparison

Qwen Max vs Wizardlm 13b

Qwen Max is the stronger model overall, scoring 34.7 to 31.4 on the Noometry Index.

Last verified . 10 shared benchmarks.

Qwen Max Alibaba (Qwen)

34.7

Rank #230 Confirmed

Wizardlm 13b Microsoft

31.4

Rank #274 Confirmed

Summary

  • They share 10 benchmarks with published results for both. Qwen Max scores higher in 6 categories and Wizardlm 13b in 1 category; 6 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Qwen Max leads 47.8 to 30.5.
  • Wizardlm 13b has downloadable open weights; the other is API-only.

Side by side

Qwen Max and Wizardlm 13b specifications
Qwen MaxWizardlm 13b
ProviderAlibaba (Qwen)Microsoft
Noometry Index34.731.4
Released2024-04-03—
WeightsProprietaryOpen
Context window33K—
Max output8K—
Input $ / M tokens$1.60—
Output $ / M tokens$6.40—
Results tracked2310

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

Qwen Max: 30.7 (#292), Wizardlm 13b: 30.1 (#298)

Coding benchmarks
BenchmarkQwen MaxWizardlm 13b
LMArena Coding12881035
Aider Polyglot21.8%—

Reasoning Qwen Max leads

Qwen Max: 25.1 (#151), Wizardlm 13b: 19.4 (#259)

Reasoning benchmarks
BenchmarkQwen MaxWizardlm 13b
LMArena Hard Prompts12691018

Math Wizardlm 13b leads

Qwen Max: 22.3 (#276), Wizardlm 13b: 30.2 (#238)

Math benchmarks
BenchmarkQwen MaxWizardlm 13b
LMArena Math12751017
OTIS Mock AIME 2024-202516.1%—
MATH Level 567.2%—
FrontierMath (Feb 2025 set)1%—

Knowledge Not comparable

Qwen Max: 30.3 (#228), Wizardlm 13b: —

Knowledge benchmarks
BenchmarkQwen MaxWizardlm 13b
GPQA Diamond56.1%—
LMArena Expert1248—

Multilingual Qwen Max leads

Qwen Max: 41.8 (#202), Wizardlm 13b: 27.1 (#277)

Multilingual benchmarks
BenchmarkQwen MaxWizardlm 13b
LMArena Non-English12631034
LMArena Chinese12541023
LMArena French1330—
LMArena German1254—
LMArena Japanese1205—
LMArena Korean1142—
LMArena Russian1274—
LMArena Spanish1290—

Instruction Following Qwen Max leads

Qwen Max: 66.5 (#208), Wizardlm 13b: 53.5 (#285)

Instruction Following benchmarks
BenchmarkQwen MaxWizardlm 13b
LMArena Instruction Following12621048

Long Context Qwen Max leads

Qwen Max: 39.4 (#180), Wizardlm 13b: 32.0 (#273)

Long Context benchmarks
BenchmarkQwen MaxWizardlm 13b
LMArena Longer Query12881054
Fiction.LiveBench66.7%—

Writing & Preference Qwen Max leads

Qwen Max: 47.8 (#205), Wizardlm 13b: 30.5 (#287)

Writing & Preference benchmarks
BenchmarkQwen MaxWizardlm 13b
LMArena Text12821077
LMArena Creative Writing12481091
LMArena Multi-Turn12771047

Frequently asked questions

Is Qwen Max better than Wizardlm 13b?

Qwen Max is the stronger model overall, scoring 34.7 to 31.4 on the Noometry Index.

Is Qwen Max or Wizardlm 13b better for coding?

They score almost the same on coding (30.7 vs 30.1); test both on your own repository before choosing.

How many benchmarks do Qwen Max and Wizardlm 13b share?

10 benchmarks have published results for both models. Qwen Max has 23 scored results on Noometry and Wizardlm 13b has 10.

Related comparisons

Go deeper