Model comparison

Claude 2.1 vs DeepSeek-R1-Distill-Qwen-1.5B

Claude 2.1 and DeepSeek-R1-Distill-Qwen-1.5B score almost the same on the Noometry Index (25.2 vs 26.1), so choose on price, context window or the category you care about most.

Last verified . 2 shared benchmarks.

Claude 2.1 Anthropic

25.2

Rank #345 Reported

Summary

  • They share 2 benchmarks with published results for both. Claude 2.1 scores higher in 2 categories and DeepSeek-R1-Distill-Qwen-1.5B in 2 categories; 3 gaps are clear of the uncertainty.
  • The widest gap is in math, where DeepSeek-R1-Distill-Qwen-1.5B leads 23.0 to 10.2.
  • The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 1.9% for Claude 2.1 and 21.4% for DeepSeek-R1-Distill-Qwen-1.5B.
  • DeepSeek-R1-Distill-Qwen-1.5B has downloadable open weights; the other is API-only.

Side by side

Claude 2.1 and DeepSeek-R1-Distill-Qwen-1.5B specifications
Claude 2.1DeepSeek-R1-Distill-Qwen-1.5B
ProviderAnthropicDeepSeek
Noometry Index25.226.1
Released2023-11-212025-01-20
WeightsProprietaryOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked75

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Claude 2.1 leads

Claude 2.1: 26.2 (#327), DeepSeek-R1-Distill-Qwen-1.5B: 21.8 (#336)

Coding benchmarks
BenchmarkClaude 2.1DeepSeek-R1-Distill-Qwen-1.5B
WeirdML7.1%—
BigCodeBench Instruct—7%
BigCodeBench Complete—7.9%

Reasoning Claude 2.1 leads

Claude 2.1: 21.4 (#221), DeepSeek-R1-Distill-Qwen-1.5B: 19.2 (#262)

Reasoning benchmarks
BenchmarkClaude 2.1DeepSeek-R1-Distill-Qwen-1.5B
Chess Puzzles—0%
DTBench51%—
Epoch Capabilities Index119.27—
ForecastBench54.2—

Math DeepSeek-R1-Distill-Qwen-1.5B leads

Claude 2.1: 10.2 (#315), DeepSeek-R1-Distill-Qwen-1.5B: 23.0 (#274)

Math benchmarks
BenchmarkClaude 2.1DeepSeek-R1-Distill-Qwen-1.5B
OTIS Mock AIME 2024-20251.9%21.4%

Knowledge Too close to call

Claude 2.1: 15.4 (#292), DeepSeek-R1-Distill-Qwen-1.5B: 16.0 (#290)

Knowledge benchmarks
BenchmarkClaude 2.1DeepSeek-R1-Distill-Qwen-1.5B
GPQA Diamond33%33.6%
MMLU73.5%—

Frequently asked questions

Is Claude 2.1 better than DeepSeek-R1-Distill-Qwen-1.5B?

Claude 2.1 and DeepSeek-R1-Distill-Qwen-1.5B score almost the same on the Noometry Index (25.2 vs 26.1), so choose on price, context window or the category you care about most.

Is Claude 2.1 or DeepSeek-R1-Distill-Qwen-1.5B better for coding?

Claude 2.1 scores higher on coding benchmarks: 26.2 versus 21.8 in the Noometry coding category.

How many benchmarks do Claude 2.1 and DeepSeek-R1-Distill-Qwen-1.5B share?

2 benchmarks have published results for both models. Claude 2.1 has 7 scored results on Noometry and DeepSeek-R1-Distill-Qwen-1.5B has 5.

Related comparisons

Go deeper