Model comparison

DeepSeek-R1 vs Step 3.5 Flash

DeepSeek-R1 and Step 3.5 Flash score almost the same on the Noometry Index (42.3 vs 42.3), so choose on price, context window or the category you care about most.

Last verified . 17 shared benchmarks.

DeepSeek-R1 DeepSeek

42.3

Rank #115 Confirmed

Step 3.5 Flash StepFun

42.3

Rank #116 Confirmed

Summary

  • They share 17 benchmarks with published results for both. DeepSeek-R1 scores higher in 6 categories and Step 3.5 Flash in 2 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where DeepSeek-R1 leads 44.5 to 39.6.
  • Step 3.5 Flash is cheaper at $0.10 / $0.30 per million input/output tokens, against $0.50 / $2.15 for DeepSeek-R1.
  • Step 3.5 Flash accepts more context: 256K tokens versus 164K.
  • Step 3.5 Flash has downloadable open weights; the other is API-only.

Side by side

DeepSeek-R1 and Step 3.5 Flash specifications
DeepSeek-R1Step 3.5 Flash
ProviderDeepSeekStepFun
Noometry Index42.342.3
Released2025-01-202026-01-29
WeightsProprietaryOpen
Context window164K256K
Max output64K256K
Input $ / M tokens$0.50$0.10
Output $ / M tokens$2.15$0.30
Results tracked5219

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding DeepSeek-R1 leads

DeepSeek-R1: 46.3 (#68), Step 3.5 Flash: 42.4 (#105)

Coding benchmarks
BenchmarkDeepSeek-R1Step 3.5 Flash
LMArena Coding14271436
Aider Polyglot71.4%—
SciCode35.7%—
WeirdML41.6%—
LiveBench Coding66.7%—
ALE-Bench804.12—
AlgoTune1.7—

Agentic & Tool Use Not comparable

DeepSeek-R1: 30.7 (#75), Step 3.5 Flash: —

Agentic & Tool Use benchmarks
BenchmarkDeepSeek-R1Step 3.5 Flash
DeepResearch Bench35.1%—
BALROG34.9%—
METR Time Horizons53.8%—

Reasoning Step 3.5 Flash leads

DeepSeek-R1: 18.6 (#278), Step 3.5 Flash: 22.2 (#202)

Reasoning benchmarks
BenchmarkDeepSeek-R1Step 3.5 Flash
LMArena Hard Prompts14161411
ARC-AGI-21.3%—
SimpleBench40.8%—
Kagi LLM Benchmark69.4%—
NYT Connections (extended)—28.4%
ARC-AGI-121.2%—
CritPt1.1%—
LiveBench Reasoning83.2%—
LiveBench Data Analysis69.8%—
Epoch Capabilities Index141.29—
ForecastBench60—
LiveBench71.6%—

Math DeepSeek-R1 leads

DeepSeek-R1: 43.8 (#79), Step 3.5 Flash: 42.6 (#84)

Math benchmarks
BenchmarkDeepSeek-R1Step 3.5 Flash
LMArena Math14001408
MathArena Final-Answer Competitions—66.8%
OTIS Mock AIME 2024-202566.4%—
Omni-MATH42.4%—
LiveBench Math80.7%—
MATH Level 596.6%—

Knowledge DeepSeek-R1 leads

DeepSeek-R1: 44.5 (#87), Step 3.5 Flash: 39.6 (#132)

Knowledge benchmarks
BenchmarkDeepSeek-R1Step 3.5 Flash
LMArena Expert13941421
GPQA Diamond76.3%—
MMLU-Pro79.3%—
Confabulations12.7%—
Vectara Hallucination Rate11.3%—
GPQA (HELM)66.6%—

Multilingual DeepSeek-R1 leads

DeepSeek-R1: 52.4 (#85), Step 3.5 Flash: 50.5 (#119)

Multilingual benchmarks
BenchmarkDeepSeek-R1Step 3.5 Flash
LMArena Non-English14121385
LMArena Chinese14421447
LMArena French14171421
LMArena German14041405
LMArena Japanese13911354
LMArena Korean13601352
LMArena Russian14231385
LMArena Spanish14111419

Instruction Following Step 3.5 Flash leads

DeepSeek-R1: 72.0 (#143), Step 3.5 Flash: 73.1 (#124)

Instruction Following benchmarks
BenchmarkDeepSeek-R1Step 3.5 Flash
LMArena Instruction Following13821385
LiveBench Instruction Following80.5%—
IFEval78.4%—

Long Context DeepSeek-R1 leads

DeepSeek-R1: 45.4 (#36), Step 3.5 Flash: 42.8 (#117)

Long Context benchmarks
BenchmarkDeepSeek-R1Step 3.5 Flash
LMArena Longer Query13911402
Fiction.LiveBench75%—

Writing & Preference DeepSeek-R1 leads

DeepSeek-R1: 61.4 (#88), Step 3.5 Flash: 58.8 (#113)

Writing & Preference benchmarks
BenchmarkDeepSeek-R1Step 3.5 Flash
LMArena Text14281403
LMArena Creative Writing14051357
LMArena Multi-Turn14051405
Short-Story Creative Writing83%—
EQ-Bench Creative Writing1500—
WildBench82.8%—
LiveBench Language48.5%—

Frequently asked questions

Is DeepSeek-R1 better than Step 3.5 Flash?

DeepSeek-R1 and Step 3.5 Flash score almost the same on the Noometry Index (42.3 vs 42.3), so choose on price, context window or the category you care about most.

Which is cheaper, DeepSeek-R1 or Step 3.5 Flash?

Step 3.5 Flash is cheaper. It lists at $0.10 per million input tokens and $0.30 per million output tokens; DeepSeek-R1 lists at $0.50 and $2.15.

Is DeepSeek-R1 or Step 3.5 Flash better for coding?

DeepSeek-R1 scores higher on coding benchmarks: 46.3 versus 42.4 in the Noometry coding category.

Which has the bigger context window?

Step 3.5 Flash does, with 256K tokens against 164K.

How many benchmarks do DeepSeek-R1 and Step 3.5 Flash share?

17 benchmarks have published results for both models. DeepSeek-R1 has 52 scored results on Noometry and Step 3.5 Flash has 19.

Related comparisons

Go deeper