Model comparison

GPT-6 Astra vs Grok 4.5

GPT-6 Astra is the stronger model overall, scoring 70.8 to 55.0 on the Noometry Index. Grok 4.5 costs 6.7× less per token, which makes it the better buy when GPT-6 Astra's lead doesn't matter for your workload.

Last verified . 45 shared benchmarks.

GPT-6 Astra OpenAI

70.8

Rank #1 Confirmed

Grok 4.5 xAI

55.0

Rank #25 Confirmed

Summary

  • They share 45 benchmarks with published results for both. GPT-6 Astra scores higher in 8 categories and Grok 4.5 in 2 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in math, where GPT-6 Astra leads 93.5 to 60.9.
  • The biggest single-benchmark swing is FrontierMath Tier 4: 97.6% for GPT-6 Astra and 24.4% for Grok 4.5.
  • Grok 4.5 is cheaper at $2 / $6 per million input/output tokens, against $10 / $50 for GPT-6 Astra.
  • GPT-6 Astra accepts more context: 1.05M tokens versus 500K.

Side by side

GPT-6 Astra and Grok 4.5 specifications
GPT-6 AstraGrok 4.5
ProviderOpenAIxAI
Noometry Index70.855.0
Released2026-09-032026-07-08
WeightsProprietaryProprietary
Context window1.05M500K
Max output128K500K
Input $ / M tokens$10$2
Output $ / M tokens$50$6
Results tracked5652

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding GPT-6 Astra leads

GPT-6 Astra: 73.7 (#2), Grok 4.5: 52.2 (#35)

Coding benchmarks
BenchmarkGPT-6 AstraGrok 4.5
DeepSWE74.1%53.8%
FrontierCode53.3%42.4%
LMArena WebDev17861553
SciCode56.5%54.1%
WeirdML93.6%46.4%
LMArena Coding14871474
ALE-Bench2,9511,309
FrontierSWE65.5%—
GSO79.4%—
MirrorCode46.7%—

Agentic & Tool Use GPT-6 Astra leads

GPT-6 Astra: 52.9 (#3), Grok 4.5: 44.4 (#17)

Agentic & Tool Use benchmarks
BenchmarkGPT-6 AstraGrok 4.5
APEX-Agents64.7%56.2%
GDP.pdf34.2%14%
Vending-Bench 215,5153,887
Remote Labor Index20.8%—
τ²-bench Banking—47.9%
PostTrainBench—23.4%
BALROG68.3%—
GBAEval—65.4%
LMArena Search—1213

Reasoning GPT-6 Astra leads

GPT-6 Astra: 85.1 (#1), Grok 4.5: 56.1 (#25)

Reasoning benchmarks
BenchmarkGPT-6 AstraGrok 4.5
ARC-AGI-295%52.6%
NYT Connections (extended)98.1%79.9%
ARC-AGI-198.5%87.2%
CritPt31.7%15.4%
Chess Puzzles72%36%
LMArena Hard Prompts14621462
DTBench97.3%96.5%
LMCA64.4%45.2%
Epoch Capabilities Index166.45153.92
SimpleBench—70%
Kagi LLM Benchmark—83.5%
EBR-Bench76.2%—
Mystery Game Puzzles84%—
Surface Evolver Bench—74.4%
Bench to the Future 30.14—

Math GPT-6 Astra leads

GPT-6 Astra: 93.5 (#2), Grok 4.5: 60.9 (#35)

Math benchmarks
BenchmarkGPT-6 AstraGrok 4.5
FrontierMath (Tiers 1-3)93.7%57.2%
FrontierMath Tier 497.6%24.4%
OTIS Mock AIME 2024-2025100%97.8%
ProofBench99%31%
LMArena Math14651459
FrontierMath Erdős2.9%—

Knowledge GPT-6 Astra leads

GPT-6 Astra: 75.3 (#1), Grok 4.5: 62.3 (#24)

Knowledge benchmarks
BenchmarkGPT-6 AstraGrok 4.5
GPQA Diamond95.8%93.4%
SimpleQA Verified75.6%48.3%
LMArena Expert14831466
Humanity's Last Exam54.8%—
Vectara Hallucination Rate8.7%—

Multimodal GPT-6 Astra leads

GPT-6 Astra: 55.0 (#3), Grok 4.5: 37.6 (#72)

Multimodal benchmarks
BenchmarkGPT-6 AstraGrok 4.5
LMArena Vision12811288
Blueprint-Bench 249.7%27.3%
Furniture Assembly80%22.5%
LMArena Document14681452

Multilingual Too close to call

GPT-6 Astra: 53.7 (#61), Grok 4.5: 54.4 (#42)

Multilingual benchmarks
BenchmarkGPT-6 AstraGrok 4.5
LMArena Non-English14301440
LMArena Chinese14841496
LMArena French14561456
LMArena German14401446
LMArena Japanese13791428
LMArena Korean14261404
LMArena Russian14361448
LMArena Spanish14071450

Instruction Following Too close to call

GPT-6 Astra: 76.3 (#44), Grok 4.5: 76.0 (#48)

Instruction Following benchmarks
BenchmarkGPT-6 AstraGrok 4.5
LMArena Instruction Following14501446

Long Context Too close to call

GPT-6 Astra: 44.5 (#62), Grok 4.5: 44.8 (#56)

Long Context benchmarks
BenchmarkGPT-6 AstraGrok 4.5
LMArena Longer Query14561463

Writing & Preference GPT-6 Astra leads

GPT-6 Astra: 75.3 (#7), Grok 4.5: 65.8 (#42)

Writing & Preference benchmarks
BenchmarkGPT-6 AstraGrok 4.5
LMArena Text14411448
LMArena Creative Writing14181442
EQ-Bench Creative Writing21731579
LMArena Multi-Turn14481456

Frequently asked questions

Is GPT-6 Astra better than Grok 4.5?

GPT-6 Astra is the stronger model overall, scoring 70.8 to 55.0 on the Noometry Index. Grok 4.5 costs 6.7× less per token, which makes it the better buy when GPT-6 Astra's lead doesn't matter for your workload.

Which is cheaper, GPT-6 Astra or Grok 4.5?

Grok 4.5 is cheaper. It lists at $2 per million input tokens and $6 per million output tokens; GPT-6 Astra lists at $10 and $50.

Is GPT-6 Astra or Grok 4.5 better for coding?

GPT-6 Astra scores higher on coding benchmarks: 73.7 versus 52.2 in the Noometry coding category.

Which has the bigger context window?

GPT-6 Astra does, with 1.05M tokens against 500K.

How many benchmarks do GPT-6 Astra and Grok 4.5 share?

45 benchmarks have published results for both models. GPT-6 Astra has 56 scored results on Noometry and Grok 4.5 has 52.

Related comparisons

Go deeper