Model comparison

DeepSeek-V3 vs GPT-4

DeepSeek-V3 is the stronger model overall, scoring 39.5 to 29.1 on the Noometry Index.

Last verified . 35 shared benchmarks.

DeepSeek-V3 DeepSeek

39.5

Rank #166 Confirmed

GPT-4 OpenAI

29.1

Rank #316 Confirmed

Summary

  • They share 35 benchmarks with published results for both. DeepSeek-V3 scores higher in 7 categories and GPT-4 in 1 category; 8 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where DeepSeek-V3 leads 57.4 to 34.9.
  • The biggest single-benchmark swing is MATH Level 5: 75.5% for DeepSeek-V3 and 23% for GPT-4.
  • DeepSeek-V3 is cheaper at $0.24 / $0.90 per million input/output tokens, against $30 / $60 for GPT-4.
  • DeepSeek-V3 accepts more context: 164K tokens versus 8K.
  • DeepSeek-V3 has downloadable open weights; the other is API-only.

Side by side

DeepSeek-V3 and GPT-4 specifications
DeepSeek-V3GPT-4
ProviderDeepSeekOpenAI
Noometry Index39.529.1
Released2024-12-262023-03-14
WeightsOpenProprietary
Context window164K8K
Max output164K8K
Input $ / M tokens$0.24$30
Output $ / M tokens$0.90$60
Results tracked6038

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding DeepSeek-V3 leads

DeepSeek-V3: 42.3 (#106), GPT-4: 31.6 (#283)

Coding benchmarks
BenchmarkDeepSeek-V3GPT-4
WeirdML36.1%12.4%
BigCodeBench Instruct50%46%
LMArena Coding13681254
BigCodeBench Complete62.2%57.2%
HumanEval+86.6%79.3%
Aider Polyglot55.1%—
SciCode35.8%—
LiveBench Coding70.9%—
MBPP+73%—

Agentic & Tool Use Not comparable

DeepSeek-V3: —, GPT-4: —

Agentic & Tool Use benchmarks
BenchmarkDeepSeek-V3GPT-4
METR Time Horizons49.6%36.1%

Reasoning DeepSeek-V3 leads

DeepSeek-V3: 20.5 (#236), GPT-4: 17.8 (#289)

Reasoning benchmarks
BenchmarkDeepSeek-V3GPT-4
LMArena Hard Prompts13651241
DTBench64.8%62.7%
LMCA15.5%17.1%
BIG-Bench Hard87.5%75.1%
Epoch Capabilities Index135.94125.89
ForecastBench59.157.8
HellaSwag88.9%95.3%
WinoGrande85.2%87.5%
SimpleBench27.2%—
Kagi LLM Benchmark52.3%—
CritPt0%—
Chess Puzzles—4%
LiveBench Reasoning65.8%—
Mystery Game Puzzles—12%
LiveBench Data Analysis60.9%—
LiveBench66.9%—
PIQA84.7%—

Math DeepSeek-V3 leads

DeepSeek-V3: 32.1 (#219), GPT-4: 10.8 (#309)

Math benchmarks
BenchmarkDeepSeek-V3GPT-4
OTIS Mock AIME 2024-202537.8%1.1%
LMArena Math13731269
MATH Level 575.5%23%
Omni-MATH40.3%—
LiveBench Math73.5%—
FrontierMath (Feb 2025 set)1.7%—
GSM8K—92%

Knowledge DeepSeek-V3 leads

DeepSeek-V3: 37.5 (#155), GPT-4: 18.4 (#282)

Knowledge benchmarks
BenchmarkDeepSeek-V3GPT-4
GPQA Diamond67.6%35.7%
LMArena Expert13511211
MMLU87.2%86.4%
TriviaQA82.9%84.8%
MMLU-Pro72.3%—
Confabulations26.1%—
Vectara Hallucination Rate6.1%—
GPQA (HELM)53.8%—
ARC (AI2) Challenge95.3%—

Multilingual DeepSeek-V3 leads

DeepSeek-V3: 48.5 (#143), GPT-4: 40.6 (#215)

Multilingual benchmarks
BenchmarkDeepSeek-V3GPT-4
LMArena Non-English13581246
LMArena Chinese13911242
LMArena French13851283
LMArena German13741251
LMArena Japanese13331209
LMArena Korean13191184
LMArena Russian13731251
LMArena Spanish13581261

Instruction Following DeepSeek-V3 leads

DeepSeek-V3: 72.8 (#130), GPT-4: 65.3 (#222)

Instruction Following benchmarks
BenchmarkDeepSeek-V3GPT-4
LMArena Instruction Following13451241
LiveBench Instruction Following81.5%—
IFEval83.2%—

Long Context GPT-4 leads

DeepSeek-V3: 34.0 (#253), GPT-4: 37.7 (#212)

Long Context benchmarks
BenchmarkDeepSeek-V3GPT-4
LMArena Longer Query13521244
Fiction.LiveBench50%—

Writing & Preference DeepSeek-V3 leads

DeepSeek-V3: 57.4 (#130), GPT-4: 34.9 (#268)

Writing & Preference benchmarks
BenchmarkDeepSeek-V3GPT-4
LMArena Text13751263
LMArena Creative Writing13641244
EQ-Bench Creative Writing1472752
LMArena Multi-Turn13891257
Short-Story Creative Writing77%—
WildBench83%—
LiveBench Language49.1%—

Frequently asked questions

Is DeepSeek-V3 better than GPT-4?

DeepSeek-V3 is the stronger model overall, scoring 39.5 to 29.1 on the Noometry Index.

Which is cheaper, DeepSeek-V3 or GPT-4?

DeepSeek-V3 is cheaper. It lists at $0.24 per million input tokens and $0.90 per million output tokens; GPT-4 lists at $30 and $60.

Is DeepSeek-V3 or GPT-4 better for coding?

DeepSeek-V3 scores higher on coding benchmarks: 42.3 versus 31.6 in the Noometry coding category.

Which has the bigger context window?

DeepSeek-V3 does, with 164K tokens against 8K.

How many benchmarks do DeepSeek-V3 and GPT-4 share?

35 benchmarks have published results for both models. DeepSeek-V3 has 60 scored results on Noometry and GPT-4 has 38.

Related comparisons

Go deeper