Model comparison
Qwen3.6 Plus vs Step 5 Preview
Qwen3.6 Plus and Step 5 Preview score almost the same on the Noometry Index (47.5 vs 47.9), so choose on price, context window or the category you care about most.
Last verified . 16 shared benchmarks.
Summary
- They share 16 benchmarks with published results for both. Qwen3.6 Plus scores higher in 3 categories and Step 5 Preview in 5 categories; 4 gaps are clear of the uncertainty.
- The widest gap is in knowledge, where Qwen3.6 Plus leads 56.1 to 41.2.
- The biggest single-benchmark swing is SciCode: 40.7% for Qwen3.6 Plus and 58.9% for Step 5 Preview.
- Qwen3.6 Plus is cheaper at $0.50 / $3 per million input/output tokens, against $1 / $2.70 for Step 5 Preview.
- Step 5 Preview accepts more context: 1.02M tokens versus 1M.
Side by side
| Qwen3.6 Plus | Step 5 Preview | |
|---|---|---|
| Provider | Alibaba (Qwen) | StepFun |
| Noometry Index | 47.5 | 47.9 |
| Released | 2026-03-31 | 2026-09-16 |
| Weights | Proprietary | Proprietary |
| Context window | 1M | 1.02M |
| Max output | 66K | 66K |
| Input $ / M tokens | $0.50 | $1 |
| Output $ / M tokens | $3 | $2.70 |
| Results tracked | 37 | 18 |
Sponsored placements are available on pages like this one. Advertise on Noometry
Category by category
Coding Step 5 Preview leads
Qwen3.6 Plus: 40.8 (#130), Step 5 Preview: 51.8 (#36)
| Benchmark | Qwen3.6 Plus | Step 5 Preview |
|---|---|---|
| LMArena WebDev | 1461 | 1564 |
| SciCode | 40.7% | 58.9% |
| LMArena Coding | 1467 | 1480 |
| SWE-bench Verified | 57.9% | — |
| ALE-Bench | 670.15 | — |
Agentic & Tool Use Not comparable
Qwen3.6 Plus: —, Step 5 Preview: —
| Benchmark | Qwen3.6 Plus | Step 5 Preview |
|---|---|---|
| Vending-Bench 2 | 5,115 | — |
Reasoning Step 5 Preview leads
Qwen3.6 Plus: 29.3 (#93), Step 5 Preview: 40.0 (#57)
| Benchmark | Qwen3.6 Plus | Step 5 Preview |
|---|---|---|
| CritPt | 2.9% | 20.9% |
| LMArena Hard Prompts | 1449 | 1465 |
| NYT Connections (extended) | 60.3% | — |
| Chess Puzzles | 17% | — |
| Thematic Generalization | 59.5% | — |
| Mystery Game Puzzles | 12% | — |
| DTBench | 81.9% | — |
| LMCA | 33.1% | — |
| Epoch Capabilities Index | 147.65 | — |
Math Qwen3.6 Plus leads
Qwen3.6 Plus: 51.8 (#54), Step 5 Preview: 46.1 (#72)
| Benchmark | Qwen3.6 Plus | Step 5 Preview |
|---|---|---|
| LMArena Math | 1450 | 1470 |
| FrontierMath (Tiers 1-3) | 38.2% | — |
| OTIS Mock AIME 2024-2025 | 93.3% | — |
| ProofBench | — | 42% |
| FrontierMath (Feb 2025 set) | 26.2% | — |
| FrontierMath Tier 4 (v1) | 8.3% | — |
Knowledge Qwen3.6 Plus leads
Qwen3.6 Plus: 56.1 (#45), Step 5 Preview: 41.2 (#112)
| Benchmark | Qwen3.6 Plus | Step 5 Preview |
|---|---|---|
| LMArena Expert | 1454 | 1470 |
| GPQA Diamond | 88.4% | — |
| SimpleQA Verified | 44.1% | — |
Multimodal Not comparable
Qwen3.6 Plus: —, Step 5 Preview: 41.0 (#41)
| Benchmark | Qwen3.6 Plus | Step 5 Preview |
|---|---|---|
| LMArena Vision | — | 1267 |
Multilingual Too close to call
Qwen3.6 Plus: 53.3 (#70), Step 5 Preview: 53.8 (#54)
| Benchmark | Qwen3.6 Plus | Step 5 Preview |
|---|---|---|
| LMArena Non-English | 1424 | 1432 |
| LMArena Chinese | 1477 | 1519 |
| LMArena Russian | 1434 | 1436 |
| LMArena Spanish | 1432 | 1439 |
| LMArena French | 1455 | — |
| LMArena German | 1452 | — |
| LMArena Japanese | 1389 | — |
| LMArena Korean | 1379 | — |
Instruction Following Too close to call
Qwen3.6 Plus: 75.0 (#74), Step 5 Preview: 76.0 (#50)
| Benchmark | Qwen3.6 Plus | Step 5 Preview |
|---|---|---|
| LMArena Instruction Following | 1425 | 1444 |
Long Context Too close to call
Qwen3.6 Plus: 45.2 (#49), Step 5 Preview: 44.5 (#64)
| Benchmark | Qwen3.6 Plus | Step 5 Preview |
|---|---|---|
| LMArena Longer Query | 1439 | 1455 |
| CL-bench | 20.3% | — |
Writing & Preference Too close to call
Qwen3.6 Plus: 62.2 (#82), Step 5 Preview: 62.8 (#71)
| Benchmark | Qwen3.6 Plus | Step 5 Preview |
|---|---|---|
| LMArena Text | 1437 | 1442 |
| LMArena Creative Writing | 1404 | 1410 |
| LMArena Multi-Turn | 1438 | 1447 |
Frequently asked questions
Is Qwen3.6 Plus better than Step 5 Preview?
Qwen3.6 Plus and Step 5 Preview score almost the same on the Noometry Index (47.5 vs 47.9), so choose on price, context window or the category you care about most.
Which is cheaper, Qwen3.6 Plus or Step 5 Preview?
Qwen3.6 Plus is cheaper. It lists at $0.50 per million input tokens and $3 per million output tokens; Step 5 Preview lists at $1 and $2.70.
Is Qwen3.6 Plus or Step 5 Preview better for coding?
Step 5 Preview scores higher on coding benchmarks: 51.8 versus 40.8 in the Noometry coding category.
Which has the bigger context window?
Step 5 Preview does, with 1.02M tokens against 1M.
How many benchmarks do Qwen3.6 Plus and Step 5 Preview share?
16 benchmarks have published results for both models. Qwen3.6 Plus has 37 scored results on Noometry and Step 5 Preview has 18.