Model comparison

Grok 4.3 vs o1

Grok 4.3 is the stronger model overall, scoring 43.8 to 40.9 on the Noometry Index.

Last verified . 27 shared benchmarks.

Grok 4.3 xAI

43.8

Rank #86 Confirmed

o1 OpenAI

40.9

Rank #143 Confirmed

Summary

  • They share 27 benchmarks with published results for both. Grok 4.3 scores higher in 6 categories and o1 in 4 categories; 10 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Grok 4.3 leads 52.5 to 41.5.
  • The biggest single-benchmark swing is FrontierMath (Tiers 1-3): 42.8% for Grok 4.3 and 14.7% for o1.
  • Grok 4.3 is cheaper at $1.25 / $2.50 per million input/output tokens, against $15 / $60 for o1.
  • Grok 4.3 accepts more context: 1M tokens versus 200K.

Side by side

Grok 4.3 and o1 specifications
Grok 4.3o1
ProviderxAIOpenAI
Noometry Index43.840.9
Released2026-04-172024-09-12
WeightsProprietaryProprietary
Context window1M200K
Max output30K100K
Input $ / M tokens$1.25$15
Output $ / M tokens$2.50$60
Results tracked4052

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding o1 leads

Grok 4.3: 41.6 (#121), o1: 46.1 (#70)

Coding benchmarks
BenchmarkGrok 4.3o1
WeirdML49.9%47.6%
LMArena Coding14151367
Aider Polyglot—61.7%
LMArena WebDev1357—
SciCode47.3%—
LiveBench Coding—69.7%
CadEval—56%
ALE-Bench944.17—
HumanEval+—89%
MBPP+—80.2%

Agentic & Tool Use Grok 4.3 leads

Grok 4.3: 27.7 (#99), o1: 24.6 (#117)

Agentic & Tool Use benchmarks
BenchmarkGrok 4.3o1
Cybench—10%
GDP.pdf8%—
LMArena Search1165—
METR Time Horizons—51.1%
Vending-Bench 235.26—

Reasoning Grok 4.3 leads

Grok 4.3: 35.9 (#68), o1: 27.9 (#111)

Reasoning benchmarks
BenchmarkGrok 4.3o1
Chess Puzzles25%15%
LMArena Hard Prompts13961371
DTBench90.7%74.7%
LMCA38.3%22.3%
Epoch Capabilities Index149.16141.91
SimpleBench—41.7%
NYT Connections (extended)55.2%—
ARC-AGI-1—30.7%
CritPt8%—
EnigmaEval—5.7%
LiveBench Reasoning—91.6%
LiveBench Data Analysis—65.5%
ForecastBench60.3—
LiveBench—75.7%

Math Grok 4.3 leads

Grok 4.3: 46.0 (#74), o1: 36.1 (#175)

Math benchmarks
BenchmarkGrok 4.3o1
FrontierMath (Tiers 1-3)42.8%14.7%
OTIS Mock AIME 2024-202593.3%73.3%
LMArena Math13881388
FrontierMath Tier 414.6%—
ProofBench11%—
LiveBench Math—80.3%
MATH Level 5—94.7%
FrontierMath (Feb 2025 set)—9.3%

Knowledge Grok 4.3 leads

Grok 4.3: 52.5 (#62), o1: 41.5 (#110)

Knowledge benchmarks
BenchmarkGrok 4.3o1
GPQA Diamond88.8%76.8%
SimpleQA Verified33.2%41.1%
LMArena Expert13851361
Humanity's Last Exam—8%
Confabulations—11.7%

Multimodal o1 leads

Grok 4.3: 31.6 (#104), o1: 34.2 (#93)

Multimodal benchmarks
BenchmarkGrok 4.3o1
LMArena Vision12291168
GeoBench—80%
VPCT—37%
Blueprint-Bench 20%—
SpatialViz-Bench—41.4%

Multilingual Grok 4.3 leads

Grok 4.3: 50.5 (#120), o1: 48.6 (#142)

Multilingual benchmarks
BenchmarkGrok 4.3o1
LMArena Non-English13851358
LMArena Chinese14221394
LMArena French14121344
LMArena German13951337
LMArena Japanese13791346
LMArena Korean13561396
LMArena Russian13991356
LMArena Spanish13981345

Instruction Following o1 leads

Grok 4.3: 72.1 (#140), o1: 74.8 (#86)

Instruction Following benchmarks
BenchmarkGrok 4.3o1
LMArena Instruction Following13661367
LiveBench Instruction Following—81.5%

Long Context o1 leads

Grok 4.3: 42.5 (#123), o1: 50.3 (#9)

Long Context benchmarks
BenchmarkGrok 4.3o1
LMArena Longer Query13931378
Fiction.LiveBench—83.3%

Writing & Preference Grok 4.3 leads

Grok 4.3: 58.5 (#118), o1: 55.6 (#144)

Writing & Preference benchmarks
BenchmarkGrok 4.3o1
LMArena Text13971366
LMArena Creative Writing13801348
LMArena Multi-Turn14061369
Short-Story Creative Writing—70.2%
EQ-Bench 41075—
LiveBench Language—65.4%

Frequently asked questions

Is Grok 4.3 better than o1?

Grok 4.3 is the stronger model overall, scoring 43.8 to 40.9 on the Noometry Index.

Which is cheaper, Grok 4.3 or o1?

Grok 4.3 is cheaper. It lists at $1.25 per million input tokens and $2.50 per million output tokens; o1 lists at $15 and $60.

Is Grok 4.3 or o1 better for coding?

o1 scores higher on coding benchmarks: 46.1 versus 41.6 in the Noometry coding category.

Which has the bigger context window?

Grok 4.3 does, with 1M tokens against 200K.

How many benchmarks do Grok 4.3 and o1 share?

27 benchmarks have published results for both models. Grok 4.3 has 40 scored results on Noometry and o1 has 52.

Related comparisons

Go deeper