Model comparison

Grok 4.20 (Non-Reasoning) vs Step 5 Preview

Grok 4.20 (Non-Reasoning) and Step 5 Preview score almost the same on the Noometry Index (48.6 vs 47.9), so choose on price, context window or the category you care about most.

Last verified . 16 shared benchmarks.

Grok 4.20 (Non-Reasoning) xAI

48.6

Rank #54 Confirmed

Step 5 Preview StepFun

47.9

Rank #58 Confirmed

Summary

  • They share 16 benchmarks with published results for both. Grok 4.20 (Non-Reasoning) scores higher in 6 categories and Step 5 Preview in 3 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Grok 4.20 (Non-Reasoning) leads 52.3 to 40.0.
  • The biggest single-benchmark swing is ProofBench: 14% for Grok 4.20 (Non-Reasoning) and 42% for Step 5 Preview.
  • Step 5 Preview is cheaper at $1 / $2.70 per million input/output tokens, against $1.25 / $2.50 for Grok 4.20 (Non-Reasoning).
  • Step 5 Preview accepts more context: 1.02M tokens versus 1M.

Side by side

Grok 4.20 (Non-Reasoning) and Step 5 Preview specifications
Grok 4.20 (Non-Reasoning)Step 5 Preview
ProviderxAIStepFun
Noometry Index48.647.9
Released2026-02-172026-09-16
WeightsProprietaryProprietary
Context window1M1.02M
Max output30K66K
Input $ / M tokens$1.25$1
Output $ / M tokens$2.50$2.70
Results tracked4618

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Step 5 Preview leads

Grok 4.20 (Non-Reasoning): 42.1 (#112), Step 5 Preview: 51.8 (#36)

Coding benchmarks
BenchmarkGrok 4.20 (Non-Reasoning)Step 5 Preview
LMArena WebDev13751564
LMArena Coding14591480
SciCode—58.9%
WeirdML52.3%—
ALE-Bench1,150—

Agentic & Tool Use Not comparable

Grok 4.20 (Non-Reasoning): 34.4 (#46), Step 5 Preview: —

Agentic & Tool Use benchmarks
BenchmarkGrok 4.20 (Non-Reasoning)Step 5 Preview
Terminal-Bench57.3%—
τ²-bench Banking18%—
LMArena Search1189—
Vending-Bench 24,663—

Reasoning Grok 4.20 (Non-Reasoning) leads

Grok 4.20 (Non-Reasoning): 52.3 (#32), Step 5 Preview: 40.0 (#57)

Reasoning benchmarks
BenchmarkGrok 4.20 (Non-Reasoning)Step 5 Preview
LMArena Hard Prompts14511465
ARC-AGI-265.1%—
Kagi LLM Benchmark75%—
NYT Connections (extended)85.4%—
ARC-AGI-189.5%—
CritPt—20.9%
Chess Puzzles24%—
Thematic Generalization63.8%—
DTBench90.1%—
LMCA38.7%—
Epoch Capabilities Index151.98—
ForecastBench61.4—

Math Grok 4.20 (Non-Reasoning) leads

Grok 4.20 (Non-Reasoning): 48.2 (#65), Step 5 Preview: 46.1 (#72)

Math benchmarks
BenchmarkGrok 4.20 (Non-Reasoning)Step 5 Preview
ProofBench14%42%
LMArena Math14551470
FrontierMath (Tiers 1-3)44.9%—
FrontierMath Tier 417.1%—
OTIS Mock AIME 2024-202592.2%—

Knowledge Grok 4.20 (Non-Reasoning) leads

Grok 4.20 (Non-Reasoning): 52.8 (#60), Step 5 Preview: 41.2 (#112)

Knowledge benchmarks
BenchmarkGrok 4.20 (Non-Reasoning)Step 5 Preview
LMArena Expert14391470
GPQA Diamond89.3%—
SimpleQA Verified30.2%—

Multimodal Step 5 Preview leads

Grok 4.20 (Non-Reasoning): 33.3 (#98), Step 5 Preview: 41.0 (#41)

Multimodal benchmarks
BenchmarkGrok 4.20 (Non-Reasoning)Step 5 Preview
LMArena Vision12631267
Blueprint-Bench 20%—
LMArena Document1416—

Multilingual Too close to call

Grok 4.20 (Non-Reasoning): 54.5 (#40), Step 5 Preview: 53.8 (#54)

Multilingual benchmarks
BenchmarkGrok 4.20 (Non-Reasoning)Step 5 Preview
LMArena Non-English14411432
LMArena Chinese14811519
LMArena Russian14581436
LMArena Spanish14431439
LMArena French1476—
LMArena German1465—
LMArena Japanese1449—
LMArena Korean1417—

Instruction Following Step 5 Preview leads

Grok 4.20 (Non-Reasoning): 74.8 (#83), Step 5 Preview: 76.0 (#50)

Instruction Following benchmarks
BenchmarkGrok 4.20 (Non-Reasoning)Step 5 Preview
LMArena Instruction Following14201444

Long Context Too close to call

Grok 4.20 (Non-Reasoning): 45.5 (#34), Step 5 Preview: 44.5 (#64)

Long Context benchmarks
BenchmarkGrok 4.20 (Non-Reasoning)Step 5 Preview
LMArena Longer Query14371455
CL-bench22.2%—
CL-bench Life11.9%—

Writing & Preference Grok 4.20 (Non-Reasoning) leads

Grok 4.20 (Non-Reasoning): 65.7 (#44), Step 5 Preview: 62.8 (#71)

Writing & Preference benchmarks
BenchmarkGrok 4.20 (Non-Reasoning)Step 5 Preview
LMArena Text14511442
LMArena Creative Writing14381410
LMArena Multi-Turn14561447
EQ-Bench Creative Writing1574—

Frequently asked questions

Is Grok 4.20 (Non-Reasoning) better than Step 5 Preview?

Grok 4.20 (Non-Reasoning) and Step 5 Preview score almost the same on the Noometry Index (48.6 vs 47.9), so choose on price, context window or the category you care about most.

Which is cheaper, Grok 4.20 (Non-Reasoning) or Step 5 Preview?

Step 5 Preview is cheaper. It lists at $1 per million input tokens and $2.70 per million output tokens; Grok 4.20 (Non-Reasoning) lists at $1.25 and $2.50.

Is Grok 4.20 (Non-Reasoning) or Step 5 Preview better for coding?

Step 5 Preview scores higher on coding benchmarks: 51.8 versus 42.1 in the Noometry coding category.

Which has the bigger context window?

Step 5 Preview does, with 1.02M tokens against 1M.

How many benchmarks do Grok 4.20 (Non-Reasoning) and Step 5 Preview share?

16 benchmarks have published results for both models. Grok 4.20 (Non-Reasoning) has 46 scored results on Noometry and Step 5 Preview has 18.

Related comparisons

Go deeper