Model comparison

Grok-3 mini vs o1

Grok-3 mini and o1 score almost the same on the Noometry Index (41.2 vs 40.9), so choose on price, context window or the category you care about most.

Last verified . 28 shared benchmarks.

Grok-3 mini xAI

41.2

Rank #141 Confirmed

o1 OpenAI

40.9

Rank #143 Confirmed

Summary

  • They share 28 benchmarks with published results for both. Grok-3 mini scores higher in 3 categories and o1 in 5 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where o1 leads 27.9 to 13.6.
  • The biggest single-benchmark swing is Fiction.LiveBench: 66.7% for Grok-3 mini and 83.3% for o1.

Side by side

Grok-3 mini and o1 specifications
Grok-3 minio1
ProviderxAIOpenAI
Noometry Index41.240.9
Released2025-04-092024-09-12
WeightsProprietaryProprietary
Context window—200K
Max output—100K
Input $ / M tokens—$15
Output $ / M tokens—$60
Results tracked3552

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding o1 leads

Grok-3 mini: 40.8 (#131), o1: 46.1 (#70)

Coding benchmarks
BenchmarkGrok-3 minio1
Aider Polyglot49.3%61.7%
WeirdML42.6%47.6%
LMArena Coding13791367
LiveBench Coding—69.7%
CadEval—56%
HumanEval+—89%
MBPP+—80.2%

Agentic & Tool Use Not comparable

Grok-3 mini: —, o1: 24.6 (#117)

Agentic & Tool Use benchmarks
BenchmarkGrok-3 minio1
Cybench—10%
METR Time Horizons—51.1%

Reasoning o1 leads

Grok-3 mini: 13.6 (#334), o1: 27.9 (#111)

Reasoning benchmarks
BenchmarkGrok-3 minio1
ARC-AGI-116.5%30.7%
LMArena Hard Prompts13751371
Epoch Capabilities Index140.35141.91
ARC-AGI-20.4%—
SimpleBench—41.7%
Kagi LLM Benchmark61.3%—
Chess Puzzles—15%
EnigmaEval—5.7%
LiveBench Reasoning—91.6%
DTBench—74.7%
LiveBench Data Analysis—65.5%
LMCA—22.3%
LiveBench—75.7%

Math Grok-3 mini leads

Grok-3 mini: 42.1 (#85), o1: 36.1 (#175)

Math benchmarks
BenchmarkGrok-3 minio1
OTIS Mock AIME 2024-202577.8%73.3%
LMArena Math13861388
MATH Level 590.9%94.7%
FrontierMath (Feb 2025 set)5.9%9.3%
FrontierMath (Tiers 1-3)—14.7%
Omni-MATH31.8%—
LiveBench Math—80.3%

Knowledge Grok-3 mini leads

Grok-3 mini: 46.4 (#81), o1: 41.5 (#110)

Knowledge benchmarks
BenchmarkGrok-3 minio1
GPQA Diamond76.3%76.8%
Confabulations10.8%11.7%
LMArena Expert13951361
Humanity's Last Exam—8%
SimpleQA Verified—41.1%
MMLU-Pro79.9%—
GPQA (HELM)67.5%—

Multimodal Not comparable

Grok-3 mini: —, o1: 34.2 (#93)

Multimodal benchmarks
BenchmarkGrok-3 minio1
LMArena Vision—1168
GeoBench—80%
VPCT—37%
SpatialViz-Bench—41.4%

Multilingual Too close to call

Grok-3 mini: 48.1 (#145), o1: 48.6 (#142)

Multilingual benchmarks
BenchmarkGrok-3 minio1
LMArena Non-English13521358
LMArena Chinese13871394
LMArena French13571344
LMArena German13491337
LMArena Japanese13421346
LMArena Korean13351396
LMArena Russian13531356
LMArena Spanish13811345

Instruction Following Grok-3 mini leads

Grok-3 mini: 78.5 (#9), o1: 74.8 (#86)

Instruction Following benchmarks
BenchmarkGrok-3 minio1
LMArena Instruction Following13571367
LiveBench Instruction Following—81.5%
IFEval95.1%—

Long Context o1 leads

Grok-3 mini: 41.0 (#147), o1: 50.3 (#9)

Long Context benchmarks
BenchmarkGrok-3 minio1
Fiction.LiveBench66.7%83.3%
LMArena Longer Query13721378

Writing & Preference o1 leads

Grok-3 mini: 52.5 (#169), o1: 55.6 (#144)

Writing & Preference benchmarks
BenchmarkGrok-3 minio1
LMArena Text13701366
LMArena Creative Writing13421348
Short-Story Creative Writing73.5%70.2%
LMArena Multi-Turn13551369
WildBench65.1%—
LiveBench Language—65.4%

Frequently asked questions

Is Grok-3 mini better than o1?

Grok-3 mini and o1 score almost the same on the Noometry Index (41.2 vs 40.9), so choose on price, context window or the category you care about most.

Is Grok-3 mini or o1 better for coding?

o1 scores higher on coding benchmarks: 46.1 versus 40.8 in the Noometry coding category.

How many benchmarks do Grok-3 mini and o1 share?

28 benchmarks have published results for both models. Grok-3 mini has 35 scored results on Noometry and o1 has 52.

Related comparisons

Go deeper