StepFun, open weights
Step 3
Step 3 by StepFun ranks 149th of 354 ranked models on the Noometry Index as of October 2026, with a score of 40.5. Its strongest category is multimodal, where it ranks 86th.
Last verified
Specifications
- Noometry rank
- #149 of 354
- Index score
- 40.5
- Evidence
- Confirmed 17 results
- Provider
StepFun
- Released
- Unknown
- Weights
- Open weights
- Reasoning
- Unknown
- Context window
- —
- Max output
- —
- Input price
- Not listed
- Output price
- Not listed
- Blended price
- Not listed
- Output speed
- 7 tokens/s Kagi
- Value
- Not ranked
- Knowledge cutoff
- Unknown
Category scores
Each category score combines every public result we have in that category.
- Coding 40.1
- Reasoning 28.4
- Math 37.6
- Knowledge 36.8
- Multimodal 35.5
- Multilingual 46.3
- Instruction Following 70.4
- Long Context 40.3
- Writing & Preference 54.3
| Category | Score | Rank | Results |
|---|---|---|---|
| Coding | 40.1 | #147 | 1 |
| Reasoning | 28.4 | #105 | 2 |
| Math | 37.6 | #148 | 1 |
| Knowledge | 36.8 | #164 | 1 |
| Multimodal | 35.5 | #86 | 1 |
| Multilingual | 46.3 | #159 | 1 |
| Instruction Following | 70.4 | #164 | 1 |
| Long Context | 40.3 | #157 | 1 |
| Writing & Preference | 54.3 | #151 | 3 |
Strengths and weaknesses
Categories where Step 3 places highest and lowest among the models ranked in each, with its score against that category's median.
Strongest categories
Weakest categories
| Category | Score | vs median | Rank |
|---|---|---|---|
| Multimodal | 35.5 | −3.0 | #86 of 128, top 68% |
| Instruction Following | 70.4 | −0.9 | #164 of 305, top 54% |
| Multilingual | 46.3 | −1.1 | #159 of 297, top 54% |
Closest competitors
The models ranked just above and below Step 3. When scores are this close, price and speed are often the better way to choose.
| Model | Rank | Score | Blended $/M | Speed | |
|---|---|---|---|---|---|
| Claude Sonnet 4 | #145 | 40.8 | $6 | 31 | Compare |
| Qwen2.5-Max | #146 | 40.7 | — | — | Compare |
| Nemotron 3 Nano 30B A3B | #147 | 40.6 | $0.0875 | — | Compare |
| Granite 4.2 8B | #148 | 40.5 | $0.11 | — | Compare |
| MiniMax M1 | #150 | 40.3 | $0.96 | — | Compare |
| Nvidia Llama 3.3 Nemotron Super 49b v1.5 | #151 | 40.3 | $0.40 | — | Compare |
| Mistral Medium 3.5 | #152 | 40.2 | $3 | 50 | Compare |
| Nemotron 3 Super | #153 | 40.1 | $0.17 | — | Compare |
Sponsored placements are available on pages like this one. Advertise on Noometry
Benchmark results
Every published result we track, with its source. Bold rows are the ones used for ranking; where several exist we prefer independent runs over self-reported numbers.
Coding
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Coding | 1367 | #161 of 294, top 55% | LMArena | 2026-10-08 |
Reasoning
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| Kagi LLM Benchmark | 62.3% | #38 of 99, top 39% | Kagi LLM Benchmark | ||
| LMArena Hard Prompts | 1355 | #158 of 297, top 54% | LMArena | 2026-10-08 |
Math
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Math | 1366 | #153 of 285, top 54% | LMArena | 2026-10-08 |
Knowledge
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Expert | 1333 | #164 of 273, top 61% | LMArena | 2026-10-08 |
Multimodal
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Vision | 1177 | #89 of 122, top 73% | LMArena | 2026-10-09 |
Multilingual
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Non-English | 1327 | #159 of 297, top 54% | LMArena | 2026-10-08 | |
| LMArena Chinese | 1397 | #139 of 285, top 49% | LMArena | 2026-10-08 | |
| LMArena German | 1371 | #115 of 231, top 50% | LMArena | 2026-10-08 | |
| LMArena Korean | 1269 | #142 of 213, top 67% | LMArena | 2026-10-08 | |
| LMArena Russian | 1331 | #160 of 283, top 57% | LMArena | 2026-10-08 | |
| LMArena Spanish | 1371 | #132 of 226, top 59% | LMArena | 2026-10-08 |
Instruction Following
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Instruction Following | 1332 | #157 of 298, top 53% | LMArena | 2026-10-08 |
Long Context
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Longer Query | 1326 | #168 of 291, top 58% | LMArena | 2026-10-08 |
Writing & Preference
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Text | 1350 | #159 of 297, top 54% | LMArena | 2026-10-08 | |
| LMArena Creative Writing | 1321 | #151 of 295, top 52% | LMArena | 2026-10-08 | |
| LMArena Multi-Turn | 1341 | #160 of 295, top 55% | LMArena | 2026-10-08 |
Compare Step 3
- Step 3 vs Granite 4.2 8B
- Step 3 vs MiniMax M1
- Step 3 vs Nemotron 3 Nano 30B A3B
- Step 3 vs Nvidia Llama 3.3 Nemotron Super 49b v1.5
- Step 3 vs Qwen2.5-Max
- Step 3 vs Mistral Medium 3.5
- Step 3 vs GPT-6 Astra
- Step 3 vs Claude Fable 5.1
- Step 3 vs Gemini 3.8 Flash
- Step 3 vs Kimi K3
- Step 3 vs Grok 4.6
- Step 3 vs Qwen3.8 Max
- Step 3 vs GLM-5.3
- Step 3 vs Muse Spark 1.3
Other StepFun models
Frequently asked questions
How good is Step 3?
Step 3 by StepFun ranks 149th of 354 ranked models on the Noometry Index as of October 2026, with a score of 40.5. Its strongest category is multimodal, where it ranks 86th.
Is Step 3 open source?
Yes. Step 3's weights are downloadable; check the license for commercial terms.
How fast is Step 3?
Step 3 generated about 7 output tokens per second in the Kagi LLM Benchmark's timed runs. Speed varies by provider, load and reasoning effort.
What are Step 3's strengths and weaknesses?
Relative to other ranked models, Step 3 places best in reasoning, coding, math and lowest in multimodal, instruction following, multilingual.
What is Step 3 best at?
Its best category is multimodal, where it ranks 86th on Noometry.