Model comparison

Claude Sonnet 4.5 vs DeepSeek V4 Pro

DeepSeek V4 Pro is the stronger model overall, scoring 54.3 to 44.1 on the Noometry Index.

Last verified . 42 shared benchmarks.

Claude Sonnet 4.5 Anthropic

44.1

Rank #81 Confirmed

DeepSeek V4 Pro DeepSeek

54.3

Rank #31 Confirmed

Summary

  • They share 42 benchmarks with published results for both. Claude Sonnet 4.5 scores higher in 3 categories and DeepSeek V4 Pro in 6 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in math, where DeepSeek V4 Pro leads 64.8 to 32.3.
  • The biggest single-benchmark swing is NYT Connections (extended): 37.3% for Claude Sonnet 4.5 and 91.3% for DeepSeek V4 Pro.
  • DeepSeek V4 Pro is cheaper at $0.66 / $1.98 per million input/output tokens, against $3 / $15 for Claude Sonnet 4.5.
  • DeepSeek V4 Pro accepts more context: 1M tokens versus 200K.
  • DeepSeek V4 Pro has downloadable open weights; the other is API-only.

Side by side

Claude Sonnet 4.5 and DeepSeek V4 Pro specifications
Claude Sonnet 4.5DeepSeek V4 Pro
ProviderAnthropicDeepSeek
Noometry Index44.154.3
Released2025-09-292026-04-24
WeightsProprietaryOpen
Context window200K1M
Max output64K393K
Input $ / M tokens$3$0.66
Output $ / M tokens$15$1.98
Results tracked7348

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding DeepSeek V4 Pro leads

Claude Sonnet 4.5: 47.3 (#61), DeepSeek V4 Pro: 52.4 (#34)

Coding benchmarks
BenchmarkClaude Sonnet 4.5DeepSeek V4 Pro
SWE-bench Verified71.3%77.6%
LMArena WebDev13931582
SciCode44.7%51%
WeirdML47.7%66.2%
LMArena Coding14891470
ALE-Bench796.151,403
FrontierCode—28.6%
SWE-bench Verified (bash only)71.4%—
SWE-bench Multilingual67%—
GSO14.7%—
AlgoTune1.52—

Agentic & Tool Use Claude Sonnet 4.5 leads

Claude Sonnet 4.5: 38.3 (#32), DeepSeek V4 Pro: 32.8 (#58)

Agentic & Tool Use benchmarks
BenchmarkClaude Sonnet 4.5DeepSeek V4 Pro
Vending-Bench 23,8393,285
Terminal-Bench46.5%—
APEX-Agents—47.3%
Berkeley Function Calling Leaderboard73.2%—
GDPval42.5%—
Remote Labor Index2.1%—
τ²-bench Airline72%—
τ²-bench Banking25.3%—
τ²-bench Retail72.4%—
τ²-bench Telecom84.9%—
Cybench60%—
DeepResearch Bench52.6%—
OSWorld62.9%—
LMArena Search1159—
METR Time Horizons67.4%—

Reasoning DeepSeek V4 Pro leads

Claude Sonnet 4.5: 26.9 (#125), DeepSeek V4 Pro: 56.5 (#24)

Reasoning benchmarks
BenchmarkClaude Sonnet 4.5DeepSeek V4 Pro
ARC-AGI-213.6%61.3%
Kagi LLM Benchmark57.9%53.5%
NYT Connections (extended)37.3%91.3%
ARC-AGI-163.7%90.5%
CritPt1.1%18%
Chess Puzzles12%47%
LMArena Hard Prompts14621461
Mystery Game Puzzles17%43%
DTBench83.2%93.9%
LMCA38.8%45.5%
Epoch Capabilities Index146.84155.31
ForecastBench61.956.1
SimpleBench54.3%—
EnigmaEval6%—
EBR-Bench2.4%—
Surface Evolver Bench—40%

Math DeepSeek V4 Pro leads

Claude Sonnet 4.5: 32.3 (#216), DeepSeek V4 Pro: 64.8 (#30)

Math benchmarks
BenchmarkClaude Sonnet 4.5DeepSeek V4 Pro
FrontierMath (Tiers 1-3)23.9%64.6%
FrontierMath Tier 42.4%26.8%
OTIS Mock AIME 2024-202577.8%98.6%
ProofBench19%50%
LMArena Math14491455
MathArena Final-Answer Competitions—76.6%
Omni-MATH55.3%—
MATH Level 597.7%—
FrontierMath (Feb 2025 set)15.2%—
FrontierMath Tier 4 (v1)4.2%—

Knowledge DeepSeek V4 Pro leads

Claude Sonnet 4.5: 48.4 (#76), DeepSeek V4 Pro: 59.5 (#31)

Knowledge benchmarks
BenchmarkClaude Sonnet 4.5DeepSeek V4 Pro
GPQA Diamond82.3%91.7%
SimpleQA Verified30.7%52.9%
Vectara Hallucination Rate12%8.6%
LMArena Expert14821464
Humanity's Last Exam13.7%—
MMLU-Pro86.9%—
GPQA (HELM)68.6%—

Multimodal Not comparable

Claude Sonnet 4.5: 34.8 (#89), DeepSeek V4 Pro: —

Multimodal benchmarks
BenchmarkClaude Sonnet 4.5DeepSeek V4 Pro
VPCT39.8%—
LMArena Document1450—

Multilingual Too close to call

Claude Sonnet 4.5: 53.4 (#69), DeepSeek V4 Pro: 54.4 (#45)

Multilingual benchmarks
BenchmarkClaude Sonnet 4.5DeepSeek V4 Pro
LMArena Non-English14251439
LMArena Chinese14591486
LMArena French14581472
LMArena German14271458
LMArena Japanese13901445
LMArena Korean14031447
LMArena Russian14371453
LMArena Spanish14571458

Instruction Following DeepSeek V4 Pro leads

Claude Sonnet 4.5: 75.0 (#78), DeepSeek V4 Pro: 76.1 (#47)

Instruction Following benchmarks
BenchmarkClaude Sonnet 4.5DeepSeek V4 Pro
LMArena Instruction Following14591448
IFEval85%—

Long Context Too close to call

Claude Sonnet 4.5: 45.2 (#46), DeepSeek V4 Pro: 45.0 (#51)

Long Context benchmarks
BenchmarkClaude Sonnet 4.5DeepSeek V4 Pro
LMArena Longer Query14761458
CL-bench Life—13.5%

Writing & Preference Claude Sonnet 4.5 leads

Claude Sonnet 4.5: 66.5 (#34), DeepSeek V4 Pro: 65.5 (#46)

Writing & Preference benchmarks
BenchmarkClaude Sonnet 4.5DeepSeek V4 Pro
LMArena Text14391451
LMArena Creative Writing14421446
EQ-Bench Creative Writing16781553
LMArena Multi-Turn14651467
WildBench85.4%—
EQ-Bench 4—1166

Frequently asked questions

Is Claude Sonnet 4.5 better than DeepSeek V4 Pro?

DeepSeek V4 Pro is the stronger model overall, scoring 54.3 to 44.1 on the Noometry Index.

Which is cheaper, Claude Sonnet 4.5 or DeepSeek V4 Pro?

DeepSeek V4 Pro is cheaper. It lists at $0.66 per million input tokens and $1.98 per million output tokens; Claude Sonnet 4.5 lists at $3 and $15.

Is Claude Sonnet 4.5 or DeepSeek V4 Pro better for coding?

DeepSeek V4 Pro scores higher on coding benchmarks: 52.4 versus 47.3 in the Noometry coding category.

Which has the bigger context window?

DeepSeek V4 Pro does, with 1M tokens against 200K.

How many benchmarks do Claude Sonnet 4.5 and DeepSeek V4 Pro share?

42 benchmarks have published results for both models. Claude Sonnet 4.5 has 73 scored results on Noometry and DeepSeek V4 Pro has 48.

Related comparisons

Go deeper