Model comparison

Claude 2 vs DeepSeek LLM 67B

Claude 2 and DeepSeek LLM 67B score almost the same on the Noometry Index (25.0 vs 24.9), so choose on price, context window or the category you care about most.

Last verified . 4 shared benchmarks.

Claude 2 Anthropic

25.0

Rank #346 Reported

DeepSeek LLM 67B DeepSeek

24.9

Rank #347 Confirmed

Summary

  • They share 4 benchmarks with published results for both. Claude 2 scores higher in 3 categories and DeepSeek LLM 67B in 0 categories; 2 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Claude 2 leads 16.9 to 7.0.
  • The biggest single-benchmark swing is GPQA Diamond: 34.7% for Claude 2 and 24.6% for DeepSeek LLM 67B.
  • DeepSeek LLM 67B has downloadable open weights; the other is API-only.

Side by side

Claude 2 and DeepSeek LLM 67B specifications
Claude 2DeepSeek LLM 67B
ProviderAnthropicDeepSeek
Noometry Index25.024.9
Released2023-07-112023-11-29
WeightsProprietaryOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked815

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Not comparable

Claude 2: —, DeepSeek LLM 67B: 31.9 (#278)

Coding benchmarks
BenchmarkClaude 2DeepSeek LLM 67B
LMArena Coding—1096
HumanEval+61.6%—

Reasoning Claude 2 leads

Claude 2: 21.7 (#216), DeepSeek LLM 67B: 16.5 (#304)

Reasoning benchmarks
BenchmarkClaude 2DeepSeek LLM 67B
Epoch Capabilities Index120.13110.5
Chess Puzzles—0%
LMArena Hard Prompts—1070
DTBench51.9%—

Math Too close to call

Claude 2: 9.3 (#320), DeepSeek LLM 67B: 8.7 (#324)

Math benchmarks
BenchmarkClaude 2DeepSeek LLM 67B
OTIS Mock AIME 2024-20252.5%0.8%
MATH Level 511.7%6.4%
LMArena Math—1108

Knowledge Claude 2 leads

Claude 2: 16.9 (#287), DeepSeek LLM 67B: 7.0 (#313)

Knowledge benchmarks
BenchmarkClaude 2DeepSeek LLM 67B
GPQA Diamond34.7%24.6%
MMLU78.5%—
TriviaQA87.5%—

Multilingual Not comparable

Claude 2: —, DeepSeek LLM 67B: 29.4 (#267)

Multilingual benchmarks
BenchmarkClaude 2DeepSeek LLM 67B
LMArena Non-English—1073
LMArena Chinese—1132

Instruction Following Not comparable

Claude 2: —, DeepSeek LLM 67B: 55.4 (#277)

Instruction Following benchmarks
BenchmarkClaude 2DeepSeek LLM 67B
LMArena Instruction Following—1079

Long Context Not comparable

Claude 2: —, DeepSeek LLM 67B: 33.1 (#265)

Long Context benchmarks
BenchmarkClaude 2DeepSeek LLM 67B
LMArena Longer Query—1092

Writing & Preference Not comparable

Claude 2: —, DeepSeek LLM 67B: 31.6 (#282)

Writing & Preference benchmarks
BenchmarkClaude 2DeepSeek LLM 67B
LMArena Text—1105
LMArena Creative Writing—1067
LMArena Multi-Turn—1082

Frequently asked questions

Is Claude 2 better than DeepSeek LLM 67B?

Claude 2 and DeepSeek LLM 67B score almost the same on the Noometry Index (25.0 vs 24.9), so choose on price, context window or the category you care about most.

How many benchmarks do Claude 2 and DeepSeek LLM 67B share?

4 benchmarks have published results for both models. Claude 2 has 8 scored results on Noometry and DeepSeek LLM 67B has 15.

Related comparisons

Go deeper