Model comparison

Claude Sonnet 4 vs DeepSeek V4 Pro

DeepSeek V4 Pro is the stronger model overall, scoring 54.3 to 40.8 on the Noometry Index.

Last verified . 32 shared benchmarks.

Claude Sonnet 4 Anthropic

40.8

Rank #145 Confirmed

DeepSeek V4 Pro DeepSeek

54.3

Rank #31 Confirmed

Summary

  • They share 32 benchmarks with published results for both. Claude Sonnet 4 scores higher in 1 category and DeepSeek V4 Pro in 8 categories; 9 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where DeepSeek V4 Pro leads 56.5 to 22.9.
  • The biggest single-benchmark swing is ARC-AGI-2: 5.9% for Claude Sonnet 4 and 61.3% for DeepSeek V4 Pro.
  • DeepSeek V4 Pro is cheaper at $0.66 / $1.98 per million input/output tokens, against $3 / $15 for Claude Sonnet 4.
  • DeepSeek V4 Pro accepts more context: 1M tokens versus 200K.
  • DeepSeek V4 Pro has downloadable open weights; the other is API-only.

Side by side

Claude Sonnet 4 and DeepSeek V4 Pro specifications
Claude Sonnet 4DeepSeek V4 Pro
ProviderAnthropicDeepSeek
Noometry Index40.854.3
Released2025-05-222026-04-24
WeightsProprietaryOpen
Context window200K1M
Max output64K393K
Input $ / M tokens$3$0.66
Output $ / M tokens$15$1.98
Results tracked5848

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding DeepSeek V4 Pro leads

Claude Sonnet 4: 43.5 (#88), DeepSeek V4 Pro: 52.4 (#34)

Coding benchmarks
BenchmarkClaude Sonnet 4DeepSeek V4 Pro
SciCode40%51%
WeirdML46.1%66.2%
LMArena Coding14141470
ALE-Bench655.351,403
SWE-bench Verified—77.6%
FrontierCode—28.6%
SWE-bench Verified (bash only)64.9%—
Aider Polyglot61.3%—
LMArena WebDev—1582
GSO4.9%—

Agentic & Tool Use Claude Sonnet 4 leads

Claude Sonnet 4: 38.5 (#31), DeepSeek V4 Pro: 32.8 (#58)

Agentic & Tool Use benchmarks
BenchmarkClaude Sonnet 4DeepSeek V4 Pro
APEX-Agents—47.3%
TheAgentCompany33.1%—
Cybench35%—
DeepResearch Bench46.6%—
OSWorld43.9%—
METR Time Horizons62%—
Vending-Bench 2—3,285

Reasoning DeepSeek V4 Pro leads

Claude Sonnet 4: 22.9 (#187), DeepSeek V4 Pro: 56.5 (#24)

Reasoning benchmarks
BenchmarkClaude Sonnet 4DeepSeek V4 Pro
ARC-AGI-25.9%61.3%
Kagi LLM Benchmark73%53.5%
ARC-AGI-140%90.5%
CritPt0.3%18%
LMArena Hard Prompts13721461
DTBench77.1%93.9%
LMCA29%45.5%
Epoch Capabilities Index141.69155.31
ForecastBench60.256.1
SimpleBench45.5%—
NYT Connections (extended)—91.3%
Chess Puzzles—47%
EnigmaEval3.1%—
Mystery Game Puzzles—43%
Surface Evolver Bench—40%

Math DeepSeek V4 Pro leads

Claude Sonnet 4: 43.3 (#80), DeepSeek V4 Pro: 64.8 (#30)

Knowledge DeepSeek V4 Pro leads

Claude Sonnet 4: 41.8 (#108), DeepSeek V4 Pro: 59.5 (#31)

Knowledge benchmarks
BenchmarkClaude Sonnet 4DeepSeek V4 Pro
GPQA Diamond79.2%91.7%
Vectara Hallucination Rate10.3%8.6%
LMArena Expert13721464
Humanity's Last Exam7.8%—
SimpleQA Verified—52.9%
MMLU-Pro84.3%—
Confabulations13.2%—
GPQA (HELM)70.6%—

Multimodal Not comparable

Claude Sonnet 4: 26.2 (#121), DeepSeek V4 Pro: —

Multimodal benchmarks
BenchmarkClaude Sonnet 4DeepSeek V4 Pro
LMArena Vision1191—
GeoBench37%—
VPCT34%—
MindCube44.8%—

Multilingual DeepSeek V4 Pro leads

Claude Sonnet 4: 46.7 (#156), DeepSeek V4 Pro: 54.4 (#45)

Multilingual benchmarks
BenchmarkClaude Sonnet 4DeepSeek V4 Pro
LMArena Non-English13331439
LMArena Chinese13501486
LMArena French13631472
LMArena German13311458
LMArena Japanese13021445
LMArena Korean12911447
LMArena Russian13551453
LMArena Spanish13571458

Instruction Following DeepSeek V4 Pro leads

Claude Sonnet 4: 71.7 (#145), DeepSeek V4 Pro: 76.1 (#47)

Instruction Following benchmarks
BenchmarkClaude Sonnet 4DeepSeek V4 Pro
LMArena Instruction Following13761448
IFEval84%—

Long Context DeepSeek V4 Pro leads

Claude Sonnet 4: 33.7 (#259), DeepSeek V4 Pro: 45.0 (#51)

Long Context benchmarks
BenchmarkClaude Sonnet 4DeepSeek V4 Pro
LMArena Longer Query13981458
Fiction.LiveBench46.9%—
CL-bench Life—13.5%

Writing & Preference DeepSeek V4 Pro leads

Claude Sonnet 4: 57.1 (#132), DeepSeek V4 Pro: 65.5 (#46)

Writing & Preference benchmarks
BenchmarkClaude Sonnet 4DeepSeek V4 Pro
LMArena Text13511451
LMArena Creative Writing13451446
EQ-Bench Creative Writing14831553
LMArena Multi-Turn13761467
Short-Story Creative Writing81.4%—
WildBench83.8%—
EQ-Bench 4—1166

Frequently asked questions

Is Claude Sonnet 4 better than DeepSeek V4 Pro?

DeepSeek V4 Pro is the stronger model overall, scoring 54.3 to 40.8 on the Noometry Index.

Which is cheaper, Claude Sonnet 4 or DeepSeek V4 Pro?

DeepSeek V4 Pro is cheaper. It lists at $0.66 per million input tokens and $1.98 per million output tokens; Claude Sonnet 4 lists at $3 and $15.

Is Claude Sonnet 4 or DeepSeek V4 Pro better for coding?

DeepSeek V4 Pro scores higher on coding benchmarks: 52.4 versus 43.5 in the Noometry coding category.

Which has the bigger context window?

DeepSeek V4 Pro does, with 1M tokens against 200K.

How many benchmarks do Claude Sonnet 4 and DeepSeek V4 Pro share?

32 benchmarks have published results for both models. Claude Sonnet 4 has 58 scored results on Noometry and DeepSeek V4 Pro has 48.

Related comparisons

Go deeper