Model comparison

Grok 3 vs Step 1o Turbo 202506

Grok 3 and Step 1o Turbo 202506 score almost the same on the Noometry Index (39.9 vs 39.7), so choose on price, context window or the category you care about most.

Last verified . 13 shared benchmarks.

Grok 3 xAI

39.9

Rank #157 Confirmed

Step 1o Turbo 202506 StepFun

39.7

Rank #160 Confirmed

Summary

  • They share 13 benchmarks with published results for both. Grok 3 scores higher in 6 categories and Step 1o Turbo 202506 in 2 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Step 1o Turbo 202506 leads 26.8 to 13.7.

Side by side

Grok 3 and Step 1o Turbo 202506 specifications
Grok 3Step 1o Turbo 202506
ProviderxAIStepFun
Noometry Index39.939.7
Released2025-04-09—
WeightsProprietaryProprietary
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked4014

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Grok 3 leads

Grok 3: 41.9 (#115), Step 1o Turbo 202506: 39.2 (#160)

Coding benchmarks
BenchmarkGrok 3Step 1o Turbo 202506
LMArena Coding14321339
Aider Polyglot53.3%—
WeirdML37.2%—

Agentic & Tool Use Not comparable

Grok 3: 30.5 (#76), Step 1o Turbo 202506: —

Agentic & Tool Use benchmarks
BenchmarkGrok 3Step 1o Turbo 202506
BALROG29.5%—

Reasoning Step 1o Turbo 202506 leads

Grok 3: 13.7 (#333), Step 1o Turbo 202506: 26.8 (#129)

Reasoning benchmarks
BenchmarkGrok 3Step 1o Turbo 202506
LMArena Hard Prompts14341335
ARC-AGI-20%—
SimpleBench36.1%—
Kagi LLM Benchmark61.3%—
ARC-AGI-15.5%—
Epoch Capabilities Index138.33—

Math Grok 3 leads

Grok 3: 38.0 (#145), Step 1o Turbo 202506: 36.6 (#164)

Math benchmarks
BenchmarkGrok 3Step 1o Turbo 202506
LMArena Math13911318
OTIS Mock AIME 2024-202555.6%—
Omni-MATH46.4%—
MATH Level 588.7%—
FrontierMath (Feb 2025 set)3.8%—
FrontierMath Tier 4 (v1)0%—

Knowledge Grok 3 leads

Grok 3: 46.2 (#82), Step 1o Turbo 202506: 36.1 (#176)

Knowledge benchmarks
BenchmarkGrok 3Step 1o Turbo 202506
LMArena Expert14211308
GPQA Diamond75.8%—
MMLU-Pro78.8%—
Confabulations14.2%—
Vectara Hallucination Rate5.8%—
GPQA (HELM)65%—

Multimodal Not comparable

Grok 3: —, Step 1o Turbo 202506: 36.1 (#80)

Multimodal benchmarks
BenchmarkGrok 3Step 1o Turbo 202506
LMArena Vision—1186

Multilingual Grok 3 leads

Grok 3: 52.3 (#87), Step 1o Turbo 202506: 45.3 (#173)

Multilingual benchmarks
BenchmarkGrok 3Step 1o Turbo 202506
LMArena Non-English14101313
LMArena Chinese14481380
LMArena German14311314
LMArena Russian14161327
LMArena French1460—
LMArena Japanese1387—
LMArena Korean1373—
LMArena Spanish1417—

Instruction Following Grok 3 leads

Grok 3: 75.0 (#73), Step 1o Turbo 202506: 69.2 (#176)

Instruction Following benchmarks
BenchmarkGrok 3Step 1o Turbo 202506
LMArena Instruction Following14091310
IFEval88.4%—

Long Context Step 1o Turbo 202506 leads

Grok 3: 38.7 (#192), Step 1o Turbo 202506: 41.0 (#146)

Long Context benchmarks
BenchmarkGrok 3Step 1o Turbo 202506
LMArena Longer Query14391348
Fiction.LiveBench58.3%—

Writing & Preference Grok 3 leads

Grok 3: 55.8 (#141), Step 1o Turbo 202506: 53.1 (#160)

Writing & Preference benchmarks
BenchmarkGrok 3Step 1o Turbo 202506
LMArena Text14261336
LMArena Creative Writing14141307
LMArena Multi-Turn14251340
Short-Story Creative Writing76.4%—
EQ-Bench Creative Writing1186—
WildBench84.9%—

Frequently asked questions

Is Grok 3 better than Step 1o Turbo 202506?

Grok 3 and Step 1o Turbo 202506 score almost the same on the Noometry Index (39.9 vs 39.7), so choose on price, context window or the category you care about most.

Is Grok 3 or Step 1o Turbo 202506 better for coding?

Grok 3 scores higher on coding benchmarks: 41.9 versus 39.2 in the Noometry coding category.

How many benchmarks do Grok 3 and Step 1o Turbo 202506 share?

13 benchmarks have published results for both models. Grok 3 has 40 scored results on Noometry and Step 1o Turbo 202506 has 14.

Related comparisons

Go deeper