Model comparison

Chatgpt 4o Latest 20250326 vs DeepSeek V4 Pro

DeepSeek V4 Pro is the stronger model overall, scoring 54.3 to 43.8 on the Noometry Index.

Last verified . 19 shared benchmarks.

Chatgpt 4o Latest 20250326 OpenAI

43.8

Rank #82 Confirmed

DeepSeek V4 Pro DeepSeek

54.3

Rank #31 Confirmed

Summary

  • They share 19 benchmarks with published results for both. Chatgpt 4o Latest 20250326 scores higher in 0 categories and DeepSeek V4 Pro in 8 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in math, where DeepSeek V4 Pro leads 64.8 to 38.6.
  • The biggest single-benchmark swing is Kagi LLM Benchmark: 75% for Chatgpt 4o Latest 20250326 and 53.5% for DeepSeek V4 Pro.
  • DeepSeek V4 Pro has downloadable open weights; the other is API-only.

Side by side

Chatgpt 4o Latest 20250326 and DeepSeek V4 Pro specifications
Chatgpt 4o Latest 20250326DeepSeek V4 Pro
ProviderOpenAIDeepSeek
Noometry Index43.854.3
Released—2026-04-24
WeightsProprietaryOpen
Context window—1M
Max output—393K
Input $ / M tokens—$0.66
Output $ / M tokens—$1.98
Results tracked2148

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding DeepSeek V4 Pro leads

Chatgpt 4o Latest 20250326: 41.6 (#122), DeepSeek V4 Pro: 52.4 (#34)

Coding benchmarks
BenchmarkChatgpt 4o Latest 20250326DeepSeek V4 Pro
LMArena Coding14131470
SWE-bench Verified—77.6%
FrontierCode—28.6%
LMArena WebDev—1582
SciCode—51%
WeirdML—66.2%
ALE-Bench—1,403

Agentic & Tool Use Not comparable

Chatgpt 4o Latest 20250326: —, DeepSeek V4 Pro: 32.8 (#58)

Agentic & Tool Use benchmarks
BenchmarkChatgpt 4o Latest 20250326DeepSeek V4 Pro
APEX-Agents—47.3%
Vending-Bench 2—3,285

Reasoning DeepSeek V4 Pro leads

Chatgpt 4o Latest 20250326: 33.8 (#71), DeepSeek V4 Pro: 56.5 (#24)

Reasoning benchmarks
BenchmarkChatgpt 4o Latest 20250326DeepSeek V4 Pro
Kagi LLM Benchmark75%53.5%
LMArena Hard Prompts14241461
ARC-AGI-2—61.3%
NYT Connections (extended)—91.3%
ARC-AGI-1—90.5%
CritPt—18%
Chess Puzzles—47%
Mystery Game Puzzles—43%
DTBench—93.9%
LMCA—45.5%
Surface Evolver Bench—40%
Epoch Capabilities Index—155.31
ForecastBench—56.1

Math DeepSeek V4 Pro leads

Chatgpt 4o Latest 20250326: 38.6 (#134), DeepSeek V4 Pro: 64.8 (#30)

Math benchmarks
BenchmarkChatgpt 4o Latest 20250326DeepSeek V4 Pro
LMArena Math14071455
FrontierMath (Tiers 1-3)—64.6%
FrontierMath Tier 4—26.8%
MathArena Final-Answer Competitions—76.6%
OTIS Mock AIME 2024-2025—98.6%
ProofBench—50%

Knowledge DeepSeek V4 Pro leads

Chatgpt 4o Latest 20250326: 39.2 (#137), DeepSeek V4 Pro: 59.5 (#31)

Knowledge benchmarks
BenchmarkChatgpt 4o Latest 20250326DeepSeek V4 Pro
LMArena Expert14011464
GPQA Diamond—91.7%
SimpleQA Verified—52.9%
Confabulations16.6%—
Vectara Hallucination Rate—8.6%

Multimodal Not comparable

Chatgpt 4o Latest 20250326: 39.6 (#58), DeepSeek V4 Pro: —

Multimodal benchmarks
BenchmarkChatgpt 4o Latest 20250326DeepSeek V4 Pro
LMArena Vision1243—

Multilingual DeepSeek V4 Pro leads

Chatgpt 4o Latest 20250326: 52.9 (#76), DeepSeek V4 Pro: 54.4 (#45)

Multilingual benchmarks
BenchmarkChatgpt 4o Latest 20250326DeepSeek V4 Pro
LMArena Non-English14191439
LMArena Chinese14571486
LMArena French14461472
LMArena German14231458
LMArena Japanese14051445
LMArena Korean13961447
LMArena Russian14291453
LMArena Spanish14341458

Instruction Following DeepSeek V4 Pro leads

Chatgpt 4o Latest 20250326: 74.0 (#107), DeepSeek V4 Pro: 76.1 (#47)

Instruction Following benchmarks
BenchmarkChatgpt 4o Latest 20250326DeepSeek V4 Pro
LMArena Instruction Following14031448

Long Context DeepSeek V4 Pro leads

Chatgpt 4o Latest 20250326: 43.1 (#107), DeepSeek V4 Pro: 45.0 (#51)

Long Context benchmarks
BenchmarkChatgpt 4o Latest 20250326DeepSeek V4 Pro
LMArena Longer Query14131458
CL-bench Life—13.5%

Writing & Preference DeepSeek V4 Pro leads

Chatgpt 4o Latest 20250326: 62.6 (#72), DeepSeek V4 Pro: 65.5 (#46)

Writing & Preference benchmarks
BenchmarkChatgpt 4o Latest 20250326DeepSeek V4 Pro
LMArena Text14291451
LMArena Creative Writing14051446
EQ-Bench Creative Writing15011553
LMArena Multi-Turn14541467
EQ-Bench 4—1166

Frequently asked questions

Is Chatgpt 4o Latest 20250326 better than DeepSeek V4 Pro?

DeepSeek V4 Pro is the stronger model overall, scoring 54.3 to 43.8 on the Noometry Index.

Is Chatgpt 4o Latest 20250326 or DeepSeek V4 Pro better for coding?

DeepSeek V4 Pro scores higher on coding benchmarks: 52.4 versus 41.6 in the Noometry coding category.

How many benchmarks do Chatgpt 4o Latest 20250326 and DeepSeek V4 Pro share?

19 benchmarks have published results for both models. Chatgpt 4o Latest 20250326 has 21 scored results on Noometry and DeepSeek V4 Pro has 48.

Related comparisons

Go deeper