Model comparison

Deepseek Coder v2 vs o1

o1 is the stronger model overall, scoring 40.9 to 35.9 on the Noometry Index.

Last verified . 19 shared benchmarks.

Deepseek Coder v2 DeepSeek

35.9

Rank #220 Confirmed

o1 OpenAI

40.9

Rank #143 Confirmed

Summary

  • They share 19 benchmarks with published results for both. Deepseek Coder v2 scores higher in 0 categories and o1 in 8 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where o1 leads 55.6 to 38.2.
  • Deepseek Coder v2 has downloadable open weights; the other is API-only.

Side by side

Deepseek Coder v2 and o1 specifications
Deepseek Coder v2o1
ProviderDeepSeekOpenAI
Noometry Index35.940.9
Released2024-06-172024-09-12
WeightsOpenProprietary
Context window—200K
Max output—100K
Input $ / M tokens—$15
Output $ / M tokens—$60
Results tracked2452

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding o1 leads

Deepseek Coder v2: 38.1 (#183), o1: 46.1 (#70)

Coding benchmarks
BenchmarkDeepseek Coder v2o1
LMArena Coding12511367
HumanEval+82.3%89%
MBPP+75.1%80.2%
Aider Polyglot—61.7%
WeirdML—47.6%
BigCodeBench Instruct48.2%—
LiveBench Coding—69.7%
BigCodeBench Complete59.7%—
CadEval—56%

Agentic & Tool Use Not comparable

Deepseek Coder v2: —, o1: 24.6 (#117)

Agentic & Tool Use benchmarks
BenchmarkDeepseek Coder v2o1
Cybench—10%
METR Time Horizons—51.1%

Reasoning o1 leads

Deepseek Coder v2: 23.6 (#176), o1: 27.9 (#111)

Reasoning benchmarks
BenchmarkDeepseek Coder v2o1
LMArena Hard Prompts12071371
SimpleBench—41.7%
ARC-AGI-1—30.7%
Chess Puzzles—15%
EnigmaEval—5.7%
LiveBench Reasoning—91.6%
DTBench—74.7%
LiveBench Data Analysis—65.5%
LMCA—22.3%
Epoch Capabilities Index—141.91
LiveBench—75.7%
WinoGrande83.7%—

Math o1 leads

Deepseek Coder v2: 34.9 (#190), o1: 36.1 (#175)

Math benchmarks
BenchmarkDeepseek Coder v2o1
LMArena Math12411388
FrontierMath (Tiers 1-3)—14.7%
OTIS Mock AIME 2024-2025—73.3%
LiveBench Math—80.3%
MATH Level 5—94.7%
FrontierMath (Feb 2025 set)—9.3%
GSM8K94.5%—

Knowledge o1 leads

Deepseek Coder v2: 32.3 (#212), o1: 41.5 (#110)

Knowledge benchmarks
BenchmarkDeepseek Coder v2o1
LMArena Expert11811361
GPQA Diamond—76.8%
Humanity's Last Exam—8%
SimpleQA Verified—41.1%
Confabulations—11.7%
ARC (AI2) Challenge64.3%—

Multimodal Not comparable

Deepseek Coder v2: —, o1: 34.2 (#93)

Multimodal benchmarks
BenchmarkDeepseek Coder v2o1
LMArena Vision—1168
GeoBench—80%
VPCT—37%
SpatialViz-Bench—41.4%

Multilingual o1 leads

Deepseek Coder v2: 36.3 (#240), o1: 48.6 (#142)

Multilingual benchmarks
BenchmarkDeepseek Coder v2o1
LMArena Non-English11821358
LMArena Chinese12011394
LMArena French11851344
LMArena German11641337
LMArena Japanese11261346
LMArena Korean11041396
LMArena Russian11881356
LMArena Spanish11531345

Instruction Following o1 leads

Deepseek Coder v2: 61.7 (#242), o1: 74.8 (#86)

Instruction Following benchmarks
BenchmarkDeepseek Coder v2o1
LMArena Instruction Following11801367
LiveBench Instruction Following—81.5%

Long Context o1 leads

Deepseek Coder v2: 37.0 (#224), o1: 50.3 (#9)

Long Context benchmarks
BenchmarkDeepseek Coder v2o1
LMArena Longer Query12191378
Fiction.LiveBench—83.3%

Writing & Preference o1 leads

Deepseek Coder v2: 38.2 (#253), o1: 55.6 (#144)

Writing & Preference benchmarks
BenchmarkDeepseek Coder v2o1
LMArena Text11911366
LMArena Creative Writing11201348
LMArena Multi-Turn11771369
Short-Story Creative Writing—70.2%
LiveBench Language—65.4%

Frequently asked questions

Is Deepseek Coder v2 better than o1?

o1 is the stronger model overall, scoring 40.9 to 35.9 on the Noometry Index.

Is Deepseek Coder v2 or o1 better for coding?

o1 scores higher on coding benchmarks: 46.1 versus 38.1 in the Noometry coding category.

How many benchmarks do Deepseek Coder v2 and o1 share?

19 benchmarks have published results for both models. Deepseek Coder v2 has 24 scored results on Noometry and o1 has 52.

Related comparisons

Go deeper