Model comparison

DeepSeek-R1-Distill-Qwen-14B vs Wizardlm 70b

DeepSeek-R1-Distill-Qwen-14B and Wizardlm 70b score almost the same on the Noometry Index (32.7 vs 33.0), so choose on price, context window or the category you care about most.

Last verified . 0 shared benchmarks.

DeepSeek-R1-Distill-Qwen-14B DeepSeek

32.7

Rank #252 Confirmed

Wizardlm 70b Microsoft

33.0

Rank #249 Confirmed

Summary

  • The widest gap is in coding, where DeepSeek-R1-Distill-Qwen-14B leads 36.9 to 31.4.

Side by side

DeepSeek-R1-Distill-Qwen-14B and Wizardlm 70b specifications
DeepSeek-R1-Distill-Qwen-14BWizardlm 70b
ProviderDeepSeekMicrosoft
Noometry Index32.733.0
Released2025-01-20—
WeightsOpenOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked712

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding DeepSeek-R1-Distill-Qwen-14B leads

DeepSeek-R1-Distill-Qwen-14B: 36.9 (#200), Wizardlm 70b: 31.4 (#285)

Coding benchmarks
BenchmarkDeepSeek-R1-Distill-Qwen-14BWizardlm 70b
BigCodeBench Instruct38.1%—
LMArena Coding—1081
BigCodeBench Complete48.4%—

Reasoning Wizardlm 70b leads

DeepSeek-R1-Distill-Qwen-14B: 19.2 (#263), Wizardlm 70b: 20.7 (#234)

Reasoning benchmarks
BenchmarkDeepSeek-R1-Distill-Qwen-14BWizardlm 70b
Chess Puzzles1%—
LMArena Hard Prompts—1079
Epoch Capabilities Index135.43—

Math DeepSeek-R1-Distill-Qwen-14B leads

DeepSeek-R1-Distill-Qwen-14B: 35.5 (#184), Wizardlm 70b: 32.2 (#218)

Math benchmarks
BenchmarkDeepSeek-R1-Distill-Qwen-14BWizardlm 70b
OTIS Mock AIME 2024-202550.6%—
LMArena Math—1116
MATH Level 587.1%—

Knowledge Not comparable

DeepSeek-R1-Distill-Qwen-14B: 24.1 (#270), Wizardlm 70b: —

Knowledge benchmarks
BenchmarkDeepSeek-R1-Distill-Qwen-14BWizardlm 70b
GPQA Diamond44.7%—

Multilingual Not comparable

DeepSeek-R1-Distill-Qwen-14B: —, Wizardlm 70b: 29.6 (#265)

Multilingual benchmarks
BenchmarkDeepSeek-R1-Distill-Qwen-14BWizardlm 70b
LMArena Non-English—1078
LMArena Chinese—1052
LMArena German—1083
LMArena Russian—1155

Instruction Following Not comparable

DeepSeek-R1-Distill-Qwen-14B: —, Wizardlm 70b: 56.3 (#273)

Instruction Following benchmarks
BenchmarkDeepSeek-R1-Distill-Qwen-14BWizardlm 70b
LMArena Instruction Following—1093

Long Context Not comparable

DeepSeek-R1-Distill-Qwen-14B: —, Wizardlm 70b: 33.3 (#263)

Long Context benchmarks
BenchmarkDeepSeek-R1-Distill-Qwen-14BWizardlm 70b
LMArena Longer Query—1097

Writing & Preference Not comparable

DeepSeek-R1-Distill-Qwen-14B: —, Wizardlm 70b: 34.8 (#269)

Writing & Preference benchmarks
BenchmarkDeepSeek-R1-Distill-Qwen-14BWizardlm 70b
LMArena Text—1120
LMArena Creative Writing—1149
LMArena Multi-Turn—1108

Frequently asked questions

Is DeepSeek-R1-Distill-Qwen-14B better than Wizardlm 70b?

DeepSeek-R1-Distill-Qwen-14B and Wizardlm 70b score almost the same on the Noometry Index (32.7 vs 33.0), so choose on price, context window or the category you care about most.

Is DeepSeek-R1-Distill-Qwen-14B or Wizardlm 70b better for coding?

DeepSeek-R1-Distill-Qwen-14B scores higher on coding benchmarks: 36.9 versus 31.4 in the Noometry coding category.

How many benchmarks do DeepSeek-R1-Distill-Qwen-14B and Wizardlm 70b share?

0 benchmarks have published results for both models. DeepSeek-R1-Distill-Qwen-14B has 7 scored results on Noometry and Wizardlm 70b has 12.

Related comparisons

Go deeper