Model comparison

DeepSeek LLM 67B vs Mistral 7B

DeepSeek LLM 67B is the stronger model overall, scoring 24.9 to 23.0 on the Noometry Index.

Last verified . 15 shared benchmarks.

DeepSeek LLM 67B DeepSeek

24.9

Rank #347 Confirmed

Mistral 7B Mistral AI

23.0

Rank #351 Confirmed

Summary

  • They share 15 benchmarks with published results for both. DeepSeek LLM 67B scores higher in 7 categories and Mistral 7B in 1 category; 4 gaps are clear of the uncertainty.
  • The widest gap is in coding, where DeepSeek LLM 67B leads 31.9 to 26.4.
  • The biggest single-benchmark swing is GPQA Diamond: 24.6% for DeepSeek LLM 67B and 15.2% for Mistral 7B.

Side by side

DeepSeek LLM 67B and Mistral 7B specifications
DeepSeek LLM 67BMistral 7B
ProviderDeepSeekMistral AI
Noometry Index24.923.0
Released2023-11-292023-09-27
WeightsOpenOpen
Context window—8K
Max output—8K
Input $ / M tokens—$0.25
Output $ / M tokens—$0.25
Results tracked1537

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding DeepSeek LLM 67B leads

DeepSeek LLM 67B: 31.9 (#278), Mistral 7B: 26.4 (#326)

Coding benchmarks
BenchmarkDeepSeek LLM 67BMistral 7B
LMArena Coding10961082
BigCodeBench Instruct—19.5%
BigCodeBench Complete—27.3%
HumanEval+—36%
MBPP+—42.1%

Reasoning DeepSeek LLM 67B leads

DeepSeek LLM 67B: 16.5 (#304), Mistral 7B: 13.1 (#336)

Reasoning benchmarks
BenchmarkDeepSeek LLM 67BMistral 7B
Chess Puzzles0%0%
LMArena Hard Prompts10701067
Epoch Capabilities Index110.5112.21
DTBench—42.5%
Adversarial NLI—47.1%
BIG-Bench Hard—56.1%
HellaSwag—81%
PIQA—83%
WinoGrande—75.3%

Math Too close to call

DeepSeek LLM 67B: 8.7 (#324), Mistral 7B: 8.1 (#325)

Math benchmarks
BenchmarkDeepSeek LLM 67BMistral 7B
OTIS Mock AIME 2024-20250.8%0.3%
LMArena Math11081085
MATH Level 56.4%3.7%
GSM8K—54.4%

Knowledge Too close to call

DeepSeek LLM 67B: 7.0 (#313), Mistral 7B: 7.4 (#311)

Knowledge benchmarks
BenchmarkDeepSeek LLM 67BMistral 7B
GPQA Diamond24.6%15.2%
LMArena Expert—1036
ARC (AI2) Challenge—78.6%
BoolQ—87.4%
MMLU—62.5%
OpenBookQA—79.8%
TriviaQA—75.2%

Multilingual DeepSeek LLM 67B leads

DeepSeek LLM 67B: 29.4 (#267), Mistral 7B: 25.8 (#283)

Multilingual benchmarks
BenchmarkDeepSeek LLM 67BMistral 7B
LMArena Non-English10731012
LMArena Chinese11321009
LMArena French—1037
LMArena German—987
LMArena Japanese—878
LMArena Russian—1018
LMArena Spanish—1026

Instruction Following DeepSeek LLM 67B leads

DeepSeek LLM 67B: 55.4 (#277), Mistral 7B: 54.2 (#280)

Instruction Following benchmarks
BenchmarkDeepSeek LLM 67BMistral 7B
LMArena Instruction Following10791060

Long Context Too close to call

DeepSeek LLM 67B: 33.1 (#265), Mistral 7B: 32.2 (#271)

Long Context benchmarks
BenchmarkDeepSeek LLM 67BMistral 7B
LMArena Longer Query10921060

Writing & Preference Too close to call

DeepSeek LLM 67B: 31.6 (#282), Mistral 7B: 30.7 (#286)

Writing & Preference benchmarks
BenchmarkDeepSeek LLM 67BMistral 7B
LMArena Text11051090
LMArena Creative Writing10671068
LMArena Multi-Turn10821062

Frequently asked questions

Is DeepSeek LLM 67B better than Mistral 7B?

DeepSeek LLM 67B is the stronger model overall, scoring 24.9 to 23.0 on the Noometry Index.

Is DeepSeek LLM 67B or Mistral 7B better for coding?

DeepSeek LLM 67B scores higher on coding benchmarks: 31.9 versus 26.4 in the Noometry coding category.

How many benchmarks do DeepSeek LLM 67B and Mistral 7B share?

15 benchmarks have published results for both models. DeepSeek LLM 67B has 15 scored results on Noometry and Mistral 7B has 37.

Related comparisons

Go deeper