Model comparison

DeepSeek V4.1 Flash vs Step 5 Preview

DeepSeek V4.1 Flash is the stronger model overall, scoring 52.8 to 47.9 on the Noometry Index.

Last verified . 18 shared benchmarks.

DeepSeek V4.1 Flash DeepSeek

52.8

Rank #38 Confirmed

Step 5 Preview StepFun

47.9

Rank #58 Confirmed

Summary

  • They share 18 benchmarks with published results for both. DeepSeek V4.1 Flash scores higher in 8 categories and Step 5 Preview in 1 category; 8 gaps are clear of the uncertainty.
  • The widest gap is in math, where DeepSeek V4.1 Flash leads 66.7 to 46.1.
  • The biggest single-benchmark swing is ProofBench: 54% for DeepSeek V4.1 Flash and 42% for Step 5 Preview.
  • DeepSeek V4.1 Flash is cheaper at $0.15 / $0.60 per million input/output tokens, against $1 / $2.70 for Step 5 Preview.
  • Step 5 Preview accepts more context: 1.02M tokens versus 1M.
  • DeepSeek V4.1 Flash has downloadable open weights; the other is API-only.

Side by side

DeepSeek V4.1 Flash and Step 5 Preview specifications
DeepSeek V4.1 FlashStep 5 Preview
ProviderDeepSeekStepFun
Noometry Index52.847.9
Released2026-09-092026-09-16
WeightsOpenProprietary
Context window1M1.02M
Max output393K66K
Input $ / M tokens$0.15$1
Output $ / M tokens$0.60$2.70
Results tracked3718

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding DeepSeek V4.1 Flash leads

DeepSeek V4.1 Flash: 52.9 (#32), Step 5 Preview: 51.8 (#36)

Coding benchmarks
BenchmarkDeepSeek V4.1 FlashStep 5 Preview
LMArena WebDev16191564
SciCode51.9%58.9%
LMArena Coding15061480
ALE-Bench1,092—

Agentic & Tool Use Not comparable

DeepSeek V4.1 Flash: 31.2 (#69), Step 5 Preview: —

Agentic & Tool Use benchmarks
BenchmarkDeepSeek V4.1 FlashStep 5 Preview
APEX-Agents39.5%—
GDP.pdf19.8%—

Reasoning DeepSeek V4.1 Flash leads

DeepSeek V4.1 Flash: 50.2 (#36), Step 5 Preview: 40.0 (#57)

Reasoning benchmarks
BenchmarkDeepSeek V4.1 FlashStep 5 Preview
CritPt14.3%20.9%
LMArena Hard Prompts14831465
NYT Connections (extended)89.6%—
Mystery Game Puzzles43%—
DTBench89.9%—
LMCA47%—
Surface Evolver Bench46.3%—
Epoch Capabilities Index154.9—

Math DeepSeek V4.1 Flash leads

DeepSeek V4.1 Flash: 66.7 (#25), Step 5 Preview: 46.1 (#72)

Math benchmarks
BenchmarkDeepSeek V4.1 FlashStep 5 Preview
ProofBench54%42%
LMArena Math14771470
FrontierMath (Tiers 1-3)67.4%—
FrontierMath Tier 426.8%—
OTIS Mock AIME 2024-202598.3%—

Knowledge DeepSeek V4.1 Flash leads

DeepSeek V4.1 Flash: 57.9 (#38), Step 5 Preview: 41.2 (#112)

Knowledge benchmarks
BenchmarkDeepSeek V4.1 FlashStep 5 Preview
LMArena Expert15061470
GPQA Diamond89.8%—

Multimodal Step 5 Preview leads

DeepSeek V4.1 Flash: 39.1 (#61), Step 5 Preview: 41.0 (#41)

Multimodal benchmarks
BenchmarkDeepSeek V4.1 FlashStep 5 Preview
LMArena Vision12771267
Furniture Assembly34.2%—

Multilingual DeepSeek V4.1 Flash leads

DeepSeek V4.1 Flash: 55.0 (#35), Step 5 Preview: 53.8 (#54)

Multilingual benchmarks
BenchmarkDeepSeek V4.1 FlashStep 5 Preview
LMArena Non-English14481432
LMArena Chinese14971519
LMArena Russian14711436
LMArena Spanish14591439
LMArena French1452—
LMArena German1484—
LMArena Japanese1412—
LMArena Korean1452—

Instruction Following DeepSeek V4.1 Flash leads

DeepSeek V4.1 Flash: 77.3 (#26), Step 5 Preview: 76.0 (#50)

Instruction Following benchmarks
BenchmarkDeepSeek V4.1 FlashStep 5 Preview
LMArena Instruction Following14741444

Long Context Too close to call

DeepSeek V4.1 Flash: 45.2 (#47), Step 5 Preview: 44.5 (#64)

Long Context benchmarks
BenchmarkDeepSeek V4.1 FlashStep 5 Preview
LMArena Longer Query14751455

Writing & Preference DeepSeek V4.1 Flash leads

DeepSeek V4.1 Flash: 65.4 (#48), Step 5 Preview: 62.8 (#71)

Writing & Preference benchmarks
BenchmarkDeepSeek V4.1 FlashStep 5 Preview
LMArena Text14621442
LMArena Creative Writing14351410
LMArena Multi-Turn14571447
EQ-Bench Creative Writing1540—

Frequently asked questions

Is DeepSeek V4.1 Flash better than Step 5 Preview?

DeepSeek V4.1 Flash is the stronger model overall, scoring 52.8 to 47.9 on the Noometry Index.

Which is cheaper, DeepSeek V4.1 Flash or Step 5 Preview?

DeepSeek V4.1 Flash is cheaper. It lists at $0.15 per million input tokens and $0.60 per million output tokens; Step 5 Preview lists at $1 and $2.70.

Is DeepSeek V4.1 Flash or Step 5 Preview better for coding?

DeepSeek V4.1 Flash scores higher on coding benchmarks: 52.9 versus 51.8 in the Noometry coding category.

Which has the bigger context window?

Step 5 Preview does, with 1.02M tokens against 1M.

How many benchmarks do DeepSeek V4.1 Flash and Step 5 Preview share?

18 benchmarks have published results for both models. DeepSeek V4.1 Flash has 37 scored results on Noometry and Step 5 Preview has 18.

Related comparisons

Go deeper