StepFun, open weights

Step 3.5 Flash

Step 3.5 Flash by StepFun ranks 116th of 354 ranked models on the Noometry Index as of October 2026, with a score of 42.3. Its strongest category is math, where it ranks 84th. API pricing starts at $0.10 per million input tokens and $0.30 per million output tokens, with a 256K-token context window.

Last verified

Specifications

Noometry rank
#116 of 354
Index score
42.3
Evidence
Confirmed 19 results
Provider
StepFun
Released
January 29, 2026
Weights
Open weights
Reasoning
Yes
Context window
256K
Max output
256K
Input price
$0.10 / M
Output price
$0.30 / M
Blended price
$0.15 / M
Output speed
Not measured
Value
#23 of 219
Knowledge cutoff
January 2025
Input
text

Category scores

Each category score combines every public result we have in that category.

Step 3.5 Flash category scores
  1. Coding 42.4
  2. Reasoning 22.2
  3. Math 42.6
  4. Knowledge 39.6
  5. Multilingual 50.5
  6. Instruction Following 73.1
  7. Long Context 42.8
  8. Writing & Preference 58.8
Step 3.5 Flash category ranks
CategoryScoreRankResults
Coding42.4#1051
Reasoning22.2#2022
Math42.6#842
Knowledge39.6#1321
Multilingual50.5#1191
Instruction Following73.1#1241
Long Context42.8#1171
Writing & Preference58.8#1133

Strengths and weaknesses

Categories where Step 3.5 Flash places highest and lowest among the models ranked in each, with its score against that category's median.

Strongest categories

Step 3.5 Flash: strongest categories
CategoryScorevs medianRank
Math42.6+6.1#84 of 327, top 26%
Coding42.4+3.6#105 of 340, top 31%
Writing & Preference58.8+5.0#113 of 312, top 37%

Weakest categories

Step 3.5 Flash: weakest categories
CategoryScorevs medianRank
Reasoning22.2−1.4#202 of 350, top 58%
Knowledge39.6+2.3#132 of 314, top 43%
Instruction Following73.1+1.8#124 of 305, top 41%

Closest competitors

The models ranked just above and below Step 3.5 Flash. When scores are this close, price and speed are often the better way to choose.

Models ranked closest to Step 3.5 Flash
ModelRankScoreBlended $/MSpeed
Qwen3.5-Flash#11242.5$0.18—Compare
Nemotron 3 Ultra#11342.5$0.93—Compare
Hunyuan T1 20250711#11442.5——Compare
DeepSeek-R1#11542.3$0.9110Compare
Qwen3.6 27B#11742.2$1.35—Compare
Amazon Nova Experimental Chat 10 20#11842.1——Compare
Qwen3.5 122B-A10B#11942.1$1.10—Compare
Longcat Flash Chat#12042.1—69Compare

Sponsored placements are available on pages like this one. Advertise on Noometry

Benchmark results

Every published result we track, with its source. Bold rows are the ones used for ranking; where several exist we prefer independent runs over self-reported numbers.

Coding

Step 3.5 Flash Coding benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Coding1436#103 of 294, top 36%LMArena2026-10-08

Reasoning

Step 3.5 Flash Reasoning benchmark results
BenchmarkScorePositionSettingSourceDate
NYT Connections (extended)28.4%#74 of 91, top 82%Lech Mazur benchmarks
LMArena Hard Prompts1411#117 of 297, top 40%LMArena2026-10-08

Math

Step 3.5 Flash Math benchmark results
BenchmarkScorePositionSettingSourceDate
MathArena Final-Answer Competitions66.8%#19 of 29, top 66%MathArena
LMArena Math1408#113 of 285, top 40%LMArena2026-10-08

Knowledge

Step 3.5 Flash Knowledge benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Expert1421#105 of 273, top 39%LMArena2026-10-08

Multilingual

Step 3.5 Flash Multilingual benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Non-English1385#119 of 297, top 41%LMArena2026-10-08
LMArena Chinese1447#106 of 285, top 38%LMArena2026-10-08
LMArena French1421#98 of 223, top 44%LMArena2026-10-08
LMArena German1405#91 of 231, top 40%LMArena2026-10-08
LMArena Japanese1354#94 of 211, top 45%LMArena2026-10-08
LMArena Korean1352#104 of 213, top 49%LMArena2026-10-08
LMArena Russian1385#124 of 283, top 44%LMArena2026-10-08
LMArena Spanish1419#93 of 226, top 42%LMArena2026-10-08

Instruction Following

Step 3.5 Flash Instruction Following benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Instruction Following1385#119 of 298, top 40%LMArena2026-10-08

Long Context

Step 3.5 Flash Long Context benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Longer Query1402#115 of 291, top 40%LMArena2026-10-08

Writing & Preference

Step 3.5 Flash Writing & Preference benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Text1403#117 of 297, top 40%LMArena2026-10-08
LMArena Creative Writing1357#124 of 295, top 43%LMArena2026-10-08
LMArena Multi-Turn1405#116 of 295, top 40%LMArena2026-10-08

API pricing by provider

Step 3.5 Flash API prices
RouteInput $/MOutput $/MCached input $/MChecked
openrouter$0.10$0.30—2026-10-10
stepfun$0.10$0.30$0.022026-10-10

Compare Step 3.5 Flash

Other StepFun models

Frequently asked questions

How good is Step 3.5 Flash?

Step 3.5 Flash by StepFun ranks 116th of 354 ranked models on the Noometry Index as of October 2026, with a score of 42.3. Its strongest category is math, where it ranks 84th. API pricing starts at $0.10 per million input tokens and $0.30 per million output tokens, with a 256K-token context window.

How much does Step 3.5 Flash cost?

Step 3.5 Flash costs $0.10 per million input tokens and $0.30 per million output tokens on StepFun's own API, with cached input at $0.02.

What is Step 3.5 Flash's context window?

Step 3.5 Flash accepts up to 256K tokens of input and can write up to 256K tokens in one response.

Is Step 3.5 Flash open source?

Yes. Step 3.5 Flash's weights are downloadable from Hugging Face (stepfun-ai/Step-3.5-Flash); check the license for commercial terms.

What are Step 3.5 Flash's strengths and weaknesses?

Relative to other ranked models, Step 3.5 Flash places best in math, coding, writing & preference and lowest in reasoning, knowledge, instruction following.

What is Step 3.5 Flash best at?

Its best category is math, where it ranks 84th on Noometry.