Model comparison

Grok 4.3 vs Grok 4.5

Grok 4.5 is the stronger model overall, scoring 55.0 to 43.8 on the Noometry Index. Grok 4.3 costs 1.9× less per token, which makes it the better buy when Grok 4.5's lead doesn't matter for your workload.

Last verified . 38 shared benchmarks.

Grok 4.3 xAI

43.8

Rank #86 Confirmed

Grok 4.5 xAI

55.0

Rank #25 Confirmed

Summary

  • They share 38 benchmarks with published results for both. Grok 4.3 scores higher in 0 categories and Grok 4.5 in 10 categories; 10 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Grok 4.5 leads 56.1 to 35.9.
  • The biggest single-benchmark swing is Blueprint-Bench 2: 0% for Grok 4.3 and 27.3% for Grok 4.5.
  • Grok 4.3 is cheaper at $1.25 / $2.50 per million input/output tokens, against $2 / $6 for Grok 4.5.
  • Grok 4.3 accepts more context: 1M tokens versus 500K.

Side by side

Grok 4.3 and Grok 4.5 specifications
Grok 4.3Grok 4.5
ProviderxAIxAI
Noometry Index43.855.0
Released2026-04-172026-07-08
WeightsProprietaryProprietary
Context window1M500K
Max output30K500K
Input $ / M tokens$1.25$2
Output $ / M tokens$2.50$6
Results tracked4052

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Grok 4.5 leads

Grok 4.3: 41.6 (#121), Grok 4.5: 52.2 (#35)

Coding benchmarks
BenchmarkGrok 4.3Grok 4.5
LMArena WebDev13571553
SciCode47.3%54.1%
WeirdML49.9%46.4%
LMArena Coding14151474
ALE-Bench944.171,309
DeepSWE—53.8%
FrontierCode—42.4%

Agentic & Tool Use Grok 4.5 leads

Grok 4.3: 27.7 (#99), Grok 4.5: 44.4 (#17)

Agentic & Tool Use benchmarks
BenchmarkGrok 4.3Grok 4.5
GDP.pdf8%14%
LMArena Search11651213
Vending-Bench 235.263,887
APEX-Agents—56.2%
τ²-bench Banking—47.9%
PostTrainBench—23.4%
GBAEval—65.4%

Reasoning Grok 4.5 leads

Grok 4.3: 35.9 (#68), Grok 4.5: 56.1 (#25)

Reasoning benchmarks
BenchmarkGrok 4.3Grok 4.5
NYT Connections (extended)55.2%79.9%
CritPt8%15.4%
Chess Puzzles25%36%
LMArena Hard Prompts13961462
DTBench90.7%96.5%
LMCA38.3%45.2%
Epoch Capabilities Index149.16153.92
ARC-AGI-2—52.6%
SimpleBench—70%
Kagi LLM Benchmark—83.5%
ARC-AGI-1—87.2%
Surface Evolver Bench—74.4%
ForecastBench60.3—

Math Grok 4.5 leads

Grok 4.3: 46.0 (#74), Grok 4.5: 60.9 (#35)

Math benchmarks
BenchmarkGrok 4.3Grok 4.5
FrontierMath (Tiers 1-3)42.8%57.2%
FrontierMath Tier 414.6%24.4%
OTIS Mock AIME 2024-202593.3%97.8%
ProofBench11%31%
LMArena Math13881459

Knowledge Grok 4.5 leads

Grok 4.3: 52.5 (#62), Grok 4.5: 62.3 (#24)

Knowledge benchmarks
BenchmarkGrok 4.3Grok 4.5
GPQA Diamond88.8%93.4%
SimpleQA Verified33.2%48.3%
LMArena Expert13851466

Multimodal Grok 4.5 leads

Grok 4.3: 31.6 (#104), Grok 4.5: 37.6 (#72)

Multimodal benchmarks
BenchmarkGrok 4.3Grok 4.5
LMArena Vision12291288
Blueprint-Bench 20%27.3%
Furniture Assembly—22.5%
LMArena Document—1452

Multilingual Grok 4.5 leads

Grok 4.3: 50.5 (#120), Grok 4.5: 54.4 (#42)

Multilingual benchmarks
BenchmarkGrok 4.3Grok 4.5
LMArena Non-English13851440
LMArena Chinese14221496
LMArena French14121456
LMArena German13951446
LMArena Japanese13791428
LMArena Korean13561404
LMArena Russian13991448
LMArena Spanish13981450

Instruction Following Grok 4.5 leads

Grok 4.3: 72.1 (#140), Grok 4.5: 76.0 (#48)

Instruction Following benchmarks
BenchmarkGrok 4.3Grok 4.5
LMArena Instruction Following13661446

Long Context Grok 4.5 leads

Grok 4.3: 42.5 (#123), Grok 4.5: 44.8 (#56)

Long Context benchmarks
BenchmarkGrok 4.3Grok 4.5
LMArena Longer Query13931463

Writing & Preference Grok 4.5 leads

Grok 4.3: 58.5 (#118), Grok 4.5: 65.8 (#42)

Writing & Preference benchmarks
BenchmarkGrok 4.3Grok 4.5
LMArena Text13971448
LMArena Creative Writing13801442
LMArena Multi-Turn14061456
EQ-Bench Creative Writing—1579
EQ-Bench 41075—

Frequently asked questions

Is Grok 4.3 better than Grok 4.5?

Grok 4.5 is the stronger model overall, scoring 55.0 to 43.8 on the Noometry Index. Grok 4.3 costs 1.9× less per token, which makes it the better buy when Grok 4.5's lead doesn't matter for your workload.

Which is cheaper, Grok 4.3 or Grok 4.5?

Grok 4.3 is cheaper. It lists at $1.25 per million input tokens and $2.50 per million output tokens; Grok 4.5 lists at $2 and $6.

Is Grok 4.3 or Grok 4.5 better for coding?

Grok 4.5 scores higher on coding benchmarks: 52.2 versus 41.6 in the Noometry coding category.

Which has the bigger context window?

Grok 4.3 does, with 1M tokens against 500K.

How many benchmarks do Grok 4.3 and Grok 4.5 share?

38 benchmarks have published results for both models. Grok 4.3 has 40 scored results on Noometry and Grok 4.5 has 52.

Related comparisons

Go deeper