Model comparison

DeepSeek LLM 67B vs GPT-4o mini

DeepSeek LLM 67B and GPT-4o mini score almost the same on the Noometry Index (24.9 vs 25.5), so choose on price, context window or the category you care about most.

Last verified . 15 shared benchmarks.

DeepSeek LLM 67B DeepSeek

24.9

Rank #347 Confirmed

GPT-4o mini OpenAI

25.5

Rank #343 Confirmed

Summary

  • They share 15 benchmarks with published results for both. DeepSeek LLM 67B scores higher in 2 categories and GPT-4o mini in 6 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in multilingual, where GPT-4o mini leads 42.0 to 29.4.
  • The biggest single-benchmark swing is MATH Level 5: 6.4% for DeepSeek LLM 67B and 52.6% for GPT-4o mini.
  • DeepSeek LLM 67B has downloadable open weights; the other is API-only.

Side by side

DeepSeek LLM 67B and GPT-4o mini specifications
DeepSeek LLM 67BGPT-4o mini
ProviderDeepSeekOpenAI
Noometry Index24.925.5
Released2023-11-292024-07-18
WeightsOpenProprietary
Context window—128K
Max output—16K
Input $ / M tokens—$0.15
Output $ / M tokens—$0.60
Results tracked1560

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding DeepSeek LLM 67B leads

DeepSeek LLM 67B: 31.9 (#278), GPT-4o mini: 22.0 (#335)

Coding benchmarks
BenchmarkDeepSeek LLM 67BGPT-4o mini
LMArena Coding10961290
Aider Polyglot—3.6%
WeirdML—11.8%
BigCodeBench Instruct—46.1%
LiveBench Coding—43.1%
BigCodeBench Complete—57.4%
HumanEval+—83.5%
MBPP+—72.2%

Agentic & Tool Use Not comparable

DeepSeek LLM 67B: —, GPT-4o mini: 27.5 (#101)

Agentic & Tool Use benchmarks
BenchmarkDeepSeek LLM 67BGPT-4o mini
BALROG—17.4%

Reasoning DeepSeek LLM 67B leads

DeepSeek LLM 67B: 16.5 (#304), GPT-4o mini: 8.7 (#347)

Reasoning benchmarks
BenchmarkDeepSeek LLM 67BGPT-4o mini
Chess Puzzles0%0%
LMArena Hard Prompts10701267
Epoch Capabilities Index110.5126.56
ARC-AGI-2—0%
SimpleBench—10.7%
Kagi LLM Benchmark—28.8%
LiveBench Reasoning—32.8%
Mystery Game Puzzles—12%
DTBench—54.4%
LiveBench Data Analysis—50%
LMCA—10.4%
LiveBench—41.3%
PIQA—88.7%

Math GPT-4o mini leads

DeepSeek LLM 67B: 8.7 (#324), GPT-4o mini: 10.4 (#314)

Math benchmarks
BenchmarkDeepSeek LLM 67BGPT-4o mini
OTIS Mock AIME 2024-20250.8%6.9%
LMArena Math11081267
MATH Level 56.4%52.6%
FrontierMath (Tiers 1-3)—0.7%
Omni-MATH—28%
LiveBench Math—36.3%
GSM8K—91.3%

Knowledge GPT-4o mini leads

DeepSeek LLM 67B: 7.0 (#313), GPT-4o mini: 17.7 (#284)

Knowledge benchmarks
BenchmarkDeepSeek LLM 67BGPT-4o mini
GPQA Diamond24.6%37.7%
SimpleQA Verified—8.3%
MMLU-Pro—60.3%
Confabulations—37.2%
GPQA (HELM)—36.8%
LMArena Expert—1235
BoolQ—88.7%
MMLU—81.8%

Multimodal Not comparable

DeepSeek LLM 67B: —, GPT-4o mini: 25.9 (#122)

Multimodal benchmarks
BenchmarkDeepSeek LLM 67BGPT-4o mini
LMArena Vision—1066
Video-MME—64.8%
GeoBench—64%
VPCT—34%

Multilingual GPT-4o mini leads

DeepSeek LLM 67B: 29.4 (#267), GPT-4o mini: 42.0 (#199)

Multilingual benchmarks
BenchmarkDeepSeek LLM 67BGPT-4o mini
LMArena Non-English10731266
LMArena Chinese11321265
LMArena French—1297
LMArena German—1272
LMArena Japanese—1216
LMArena Korean—1195
LMArena Russian—1275
LMArena Spanish—1276

Instruction Following GPT-4o mini leads

DeepSeek LLM 67B: 55.4 (#277), GPT-4o mini: 61.9 (#239)

Instruction Following benchmarks
BenchmarkDeepSeek LLM 67BGPT-4o mini
LMArena Instruction Following10791258
LiveBench Instruction Following—56.8%
IFEval—78.2%

Long Context GPT-4o mini leads

DeepSeek LLM 67B: 33.1 (#265), GPT-4o mini: 39.1 (#186)

Long Context benchmarks
BenchmarkDeepSeek LLM 67BGPT-4o mini
LMArena Longer Query10921289

Writing & Preference GPT-4o mini leads

DeepSeek LLM 67B: 31.6 (#282), GPT-4o mini: 39.5 (#248)

Writing & Preference benchmarks
BenchmarkDeepSeek LLM 67BGPT-4o mini
LMArena Text11051286
LMArena Creative Writing10671268
LMArena Multi-Turn10821285
Short-Story Creative Writing—67.2%
EQ-Bench Creative Writing—873
WildBench—79.1%
LiveBench Language—28.6%

Frequently asked questions

Is DeepSeek LLM 67B better than GPT-4o mini?

DeepSeek LLM 67B and GPT-4o mini score almost the same on the Noometry Index (24.9 vs 25.5), so choose on price, context window or the category you care about most.

Is DeepSeek LLM 67B or GPT-4o mini better for coding?

DeepSeek LLM 67B scores higher on coding benchmarks: 31.9 versus 22.0 in the Noometry coding category.

How many benchmarks do DeepSeek LLM 67B and GPT-4o mini share?

15 benchmarks have published results for both models. DeepSeek LLM 67B has 15 scored results on Noometry and GPT-4o mini has 60.

Related comparisons

Go deeper