Model comparison
DeepSeek-V2.5 (Sep 2024) vs Step 3.7 Flash
DeepSeek-V2.5 (Sep 2024) and Step 3.7 Flash score almost the same on the Noometry Index (37.6 vs 37.3), so choose on price, context window or the category you care about most.
Last verified . 0 shared benchmarks.
Summary
- The widest gap is in coding, where Step 3.7 Flash leads 40.0 to 31.7.
Side by side
| DeepSeek-V2.5 (Sep 2024) | Step 3.7 Flash | |
|---|---|---|
| Provider | DeepSeek | StepFun |
| Noometry Index | 37.6 | 37.3 |
| Released | 2024-09-06 | 2026-05-29 |
| Weights | Open | Open |
| Context window | — | 256K |
| Max output | — | 256K |
| Input $ / M tokens | — | $0.18 |
| Output $ / M tokens | — | $1.11 |
| Results tracked | 22 | 5 |
Sponsored placements are available on pages like this one. Advertise on Noometry
Category by category
Coding Step 3.7 Flash leads
DeepSeek-V2.5 (Sep 2024): 31.7 (#281), Step 3.7 Flash: 40.0 (#150)
| Benchmark | DeepSeek-V2.5 (Sep 2024) | Step 3.7 Flash |
|---|---|---|
| Aider Polyglot | 17.8% | — |
| SciCode | — | 40% |
| BigCodeBench Instruct | 48.6% | — |
| LMArena Coding | 1309 | — |
| BigCodeBench Complete | 53.2% | — |
| ALE-Bench | — | 694.12 |
| HumanEval+ | 83.5% | — |
| MBPP+ | 74.1% | — |
Reasoning DeepSeek-V2.5 (Sep 2024) leads
DeepSeek-V2.5 (Sep 2024): 25.6 (#145), Step 3.7 Flash: 21.6 (#219)
| Benchmark | DeepSeek-V2.5 (Sep 2024) | Step 3.7 Flash |
|---|---|---|
| NYT Connections (extended) | — | 39.7% |
| CritPt | — | 2.3% |
| LMArena Hard Prompts | 1289 | — |
Math Step 3.7 Flash leads
DeepSeek-V2.5 (Sep 2024): 35.9 (#177), Step 3.7 Flash: 42.9 (#82)
| Benchmark | DeepSeek-V2.5 (Sep 2024) | Step 3.7 Flash |
|---|---|---|
| MathArena Final-Answer Competitions | — | 68.5% |
| LMArena Math | 1288 | — |
Knowledge Not comparable
DeepSeek-V2.5 (Sep 2024): 34.8 (#193), Step 3.7 Flash: —
| Benchmark | DeepSeek-V2.5 (Sep 2024) | Step 3.7 Flash |
|---|---|---|
| LMArena Expert | 1266 | — |
Multilingual Not comparable
DeepSeek-V2.5 (Sep 2024): 42.5 (#193), Step 3.7 Flash: —
| Benchmark | DeepSeek-V2.5 (Sep 2024) | Step 3.7 Flash |
|---|---|---|
| LMArena Non-English | 1273 | — |
| LMArena Chinese | 1318 | — |
| LMArena French | 1289 | — |
| LMArena German | 1258 | — |
| LMArena Japanese | 1228 | — |
| LMArena Korean | 1209 | — |
| LMArena Russian | 1289 | — |
| LMArena Spanish | 1248 | — |
Instruction Following Not comparable
DeepSeek-V2.5 (Sep 2024): 67.5 (#194), Step 3.7 Flash: —
| Benchmark | DeepSeek-V2.5 (Sep 2024) | Step 3.7 Flash |
|---|---|---|
| LMArena Instruction Following | 1280 | — |
Long Context Not comparable
DeepSeek-V2.5 (Sep 2024): 39.5 (#174), Step 3.7 Flash: —
| Benchmark | DeepSeek-V2.5 (Sep 2024) | Step 3.7 Flash |
|---|---|---|
| LMArena Longer Query | 1301 | — |
Writing & Preference Not comparable
DeepSeek-V2.5 (Sep 2024): 49.8 (#187), Step 3.7 Flash: —
| Benchmark | DeepSeek-V2.5 (Sep 2024) | Step 3.7 Flash |
|---|---|---|
| LMArena Text | 1294 | — |
| LMArena Creative Writing | 1285 | — |
| LMArena Multi-Turn | 1297 | — |
Frequently asked questions
Is DeepSeek-V2.5 (Sep 2024) better than Step 3.7 Flash?
DeepSeek-V2.5 (Sep 2024) and Step 3.7 Flash score almost the same on the Noometry Index (37.6 vs 37.3), so choose on price, context window or the category you care about most.
Is DeepSeek-V2.5 (Sep 2024) or Step 3.7 Flash better for coding?
Step 3.7 Flash scores higher on coding benchmarks: 40.0 versus 31.7 in the Noometry coding category.
How many benchmarks do DeepSeek-V2.5 (Sep 2024) and Step 3.7 Flash share?
0 benchmarks have published results for both models. DeepSeek-V2.5 (Sep 2024) has 22 scored results on Noometry and Step 3.7 Flash has 5.