Model comparison

DeepSeek LLM 67B vs Llama 3.2 1B

DeepSeek LLM 67B is the stronger model overall, scoring 24.9 to 20.1 on the Noometry Index.

Last verified . 14 shared benchmarks.

DeepSeek LLM 67B DeepSeek

24.9

Rank #347 Confirmed

Llama 3.2 1B Meta

20.1

Rank #354 Confirmed

Summary

  • They share 14 benchmarks with published results for both. DeepSeek LLM 67B scores higher in 6 categories and Llama 3.2 1B in 2 categories; 6 gaps are clear of the uncertainty.
  • The widest gap is in coding, where DeepSeek LLM 67B leads 31.9 to 21.1.

Side by side

DeepSeek LLM 67B and Llama 3.2 1B specifications
DeepSeek LLM 67BLlama 3.2 1B
ProviderDeepSeekMeta
Noometry Index24.920.1
Released2023-11-292024-09-24
WeightsOpenOpen
Context window—60K
Max output—54K
Input $ / M tokens—$0.027
Output $ / M tokens—$0.20
Results tracked1522

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding DeepSeek LLM 67B leads

DeepSeek LLM 67B: 31.9 (#278), Llama 3.2 1B: 21.1 (#338)

Coding benchmarks
BenchmarkDeepSeek LLM 67BLlama 3.2 1B
LMArena Coding10961070
BigCodeBench Instruct—8.2%
BigCodeBench Complete—11.3%

Agentic & Tool Use Not comparable

DeepSeek LLM 67B: —, Llama 3.2 1B: 14.6 (#150)

Agentic & Tool Use benchmarks
BenchmarkDeepSeek LLM 67BLlama 3.2 1B
Berkeley Function Calling Leaderboard—10.8%
BALROG—6.6%

Reasoning Too close to call

DeepSeek LLM 67B: 16.5 (#304), Llama 3.2 1B: 16.2 (#308)

Reasoning benchmarks
BenchmarkDeepSeek LLM 67BLlama 3.2 1B
Chess Puzzles0%0%
LMArena Hard Prompts10701044
Epoch Capabilities Index110.5101.99

Math Llama 3.2 1B leads

DeepSeek LLM 67B: 8.7 (#324), Llama 3.2 1B: 10.4 (#313)

Math benchmarks
BenchmarkDeepSeek LLM 67BLlama 3.2 1B
OTIS Mock AIME 2024-20250.8%0.6%
LMArena Math11081086
MATH Level 56.4%—

Knowledge Too close to call

DeepSeek LLM 67B: 7.0 (#313), Llama 3.2 1B: 7.2 (#312)

Knowledge benchmarks
BenchmarkDeepSeek LLM 67BLlama 3.2 1B
GPQA Diamond24.6%23.9%
LMArena Expert—1007

Multilingual DeepSeek LLM 67B leads

DeepSeek LLM 67B: 29.4 (#267), Llama 3.2 1B: 23.8 (#292)

Multilingual benchmarks
BenchmarkDeepSeek LLM 67BLlama 3.2 1B
LMArena Non-English1073973
LMArena Chinese1132959
LMArena German—1014
LMArena Russian—941

Instruction Following DeepSeek LLM 67B leads

DeepSeek LLM 67B: 55.4 (#277), Llama 3.2 1B: 52.4 (#290)

Instruction Following benchmarks
BenchmarkDeepSeek LLM 67BLlama 3.2 1B
LMArena Instruction Following10791031

Long Context DeepSeek LLM 67B leads

DeepSeek LLM 67B: 33.1 (#265), Llama 3.2 1B: 31.9 (#274)

Long Context benchmarks
BenchmarkDeepSeek LLM 67BLlama 3.2 1B
LMArena Longer Query10921050

Writing & Preference DeepSeek LLM 67B leads

DeepSeek LLM 67B: 31.6 (#282), Llama 3.2 1B: 21.3 (#310)

Writing & Preference benchmarks
BenchmarkDeepSeek LLM 67BLlama 3.2 1B
LMArena Text11051055
LMArena Creative Writing10671033
LMArena Multi-Turn10821030
EQ-Bench Creative Writing—200

Frequently asked questions

Is DeepSeek LLM 67B better than Llama 3.2 1B?

DeepSeek LLM 67B is the stronger model overall, scoring 24.9 to 20.1 on the Noometry Index.

Is DeepSeek LLM 67B or Llama 3.2 1B better for coding?

DeepSeek LLM 67B scores higher on coding benchmarks: 31.9 versus 21.1 in the Noometry coding category.

How many benchmarks do DeepSeek LLM 67B and Llama 3.2 1B share?

14 benchmarks have published results for both models. DeepSeek LLM 67B has 15 scored results on Noometry and Llama 3.2 1B has 22.

Related comparisons

Go deeper