Model comparison

Codestral vs DeepSeek V4 Pro

DeepSeek V4 Pro is the stronger model overall, scoring 54.3 to 30.6 on the Noometry Index. Codestral costs 2.2× less per token, which makes it the better buy when DeepSeek V4 Pro's lead doesn't matter for your workload.

Last verified . 2 shared benchmarks.

Codestral Mistral AI

30.6

Rank #290 Reported

DeepSeek V4 Pro DeepSeek

54.3

Rank #31 Confirmed

Summary

  • They share 2 benchmarks with published results for both. Codestral scores higher in 0 categories and DeepSeek V4 Pro in 2 categories; 2 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where DeepSeek V4 Pro leads 56.5 to 19.8.
  • The biggest single-benchmark swing is Kagi LLM Benchmark: 32.5% for Codestral and 53.5% for DeepSeek V4 Pro.
  • Codestral is cheaper at $0.30 / $0.90 per million input/output tokens, against $0.66 / $1.98 for DeepSeek V4 Pro.
  • DeepSeek V4 Pro accepts more context: 1M tokens versus 256K.
  • DeepSeek V4 Pro has downloadable open weights; the other is API-only.

Side by side

Codestral and DeepSeek V4 Pro specifications
CodestralDeepSeek V4 Pro
ProviderMistral AIDeepSeek
Noometry Index30.654.3
Released2024-05-292026-04-24
WeightsProprietaryOpen
Context window256K1M
Max output8K393K
Input $ / M tokens$0.30$0.66
Output $ / M tokens$0.90$1.98
Results tracked748

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding DeepSeek V4 Pro leads

Codestral: 27.3 (#321), DeepSeek V4 Pro: 52.4 (#34)

Coding benchmarks
BenchmarkCodestralDeepSeek V4 Pro
ALE-Bench137.781,403
SWE-bench Verified—77.6%
FrontierCode—28.6%
Aider Polyglot11.1%—
LMArena WebDev—1582
SciCode—51%
WeirdML—66.2%
BigCodeBench Instruct41.8%—
LMArena Coding—1470
BigCodeBench Complete52.5%—
HumanEval+73.8%—
MBPP+61.9%—

Agentic & Tool Use Not comparable

Codestral: —, DeepSeek V4 Pro: 32.8 (#58)

Agentic & Tool Use benchmarks
BenchmarkCodestralDeepSeek V4 Pro
APEX-Agents—47.3%
Vending-Bench 2—3,285

Reasoning DeepSeek V4 Pro leads

Codestral: 19.8 (#251), DeepSeek V4 Pro: 56.5 (#24)

Reasoning benchmarks
BenchmarkCodestralDeepSeek V4 Pro
Kagi LLM Benchmark32.5%53.5%
ARC-AGI-2—61.3%
NYT Connections (extended)—91.3%
ARC-AGI-1—90.5%
CritPt—18%
Chess Puzzles—47%
LMArena Hard Prompts—1461
Mystery Game Puzzles—43%
DTBench—93.9%
LMCA—45.5%
Surface Evolver Bench—40%
Epoch Capabilities Index—155.31
ForecastBench—56.1

Math Not comparable

Codestral: —, DeepSeek V4 Pro: 64.8 (#30)

Math benchmarks
BenchmarkCodestralDeepSeek V4 Pro
FrontierMath (Tiers 1-3)—64.6%
FrontierMath Tier 4—26.8%
MathArena Final-Answer Competitions—76.6%
OTIS Mock AIME 2024-2025—98.6%
ProofBench—50%
LMArena Math—1455

Knowledge Not comparable

Codestral: —, DeepSeek V4 Pro: 59.5 (#31)

Knowledge benchmarks
BenchmarkCodestralDeepSeek V4 Pro
GPQA Diamond—91.7%
SimpleQA Verified—52.9%
Vectara Hallucination Rate—8.6%
LMArena Expert—1464

Multilingual Not comparable

Codestral: —, DeepSeek V4 Pro: 54.4 (#45)

Multilingual benchmarks
BenchmarkCodestralDeepSeek V4 Pro
LMArena Non-English—1439
LMArena Chinese—1486
LMArena French—1472
LMArena German—1458
LMArena Japanese—1445
LMArena Korean—1447
LMArena Russian—1453
LMArena Spanish—1458

Instruction Following Not comparable

Codestral: —, DeepSeek V4 Pro: 76.1 (#47)

Instruction Following benchmarks
BenchmarkCodestralDeepSeek V4 Pro
LMArena Instruction Following—1448

Long Context Not comparable

Codestral: —, DeepSeek V4 Pro: 45.0 (#51)

Long Context benchmarks
BenchmarkCodestralDeepSeek V4 Pro
CL-bench Life—13.5%
LMArena Longer Query—1458

Writing & Preference Not comparable

Codestral: —, DeepSeek V4 Pro: 65.5 (#46)

Writing & Preference benchmarks
BenchmarkCodestralDeepSeek V4 Pro
LMArena Text—1451
LMArena Creative Writing—1446
EQ-Bench Creative Writing—1553
EQ-Bench 4—1166
LMArena Multi-Turn—1467

Frequently asked questions

Is Codestral better than DeepSeek V4 Pro?

DeepSeek V4 Pro is the stronger model overall, scoring 54.3 to 30.6 on the Noometry Index. Codestral costs 2.2× less per token, which makes it the better buy when DeepSeek V4 Pro's lead doesn't matter for your workload.

Which is cheaper, Codestral or DeepSeek V4 Pro?

Codestral is cheaper. It lists at $0.30 per million input tokens and $0.90 per million output tokens; DeepSeek V4 Pro lists at $0.66 and $1.98.

Is Codestral or DeepSeek V4 Pro better for coding?

DeepSeek V4 Pro scores higher on coding benchmarks: 52.4 versus 27.3 in the Noometry coding category.

Which has the bigger context window?

DeepSeek V4 Pro does, with 1M tokens against 256K.

How many benchmarks do Codestral and DeepSeek V4 Pro share?

2 benchmarks have published results for both models. Codestral has 7 scored results on Noometry and DeepSeek V4 Pro has 48.

Related comparisons

Go deeper