Model comparison

GPT-6 Luna vs Grok 4.7

GPT-6 Luna and Grok 4.7 score almost the same on the Noometry Index (53.3 vs 53.1), so choose on price, context window or the category you care about most.

Last verified . 35 shared benchmarks.

GPT-6 Luna OpenAI

53.3

Rank #36 Confirmed

Grok 4.7 xAI

53.1

Rank #37 Confirmed

Summary

  • They share 35 benchmarks with published results for both. GPT-6 Luna scores higher in 3 categories and Grok 4.7 in 7 categories; 6 gaps are clear of the uncertainty.
  • The widest gap is in math, where GPT-6 Luna leads 76.1 to 57.8.
  • The biggest single-benchmark swing is FrontierMath Tier 4: 56.1% for GPT-6 Luna and 17.1% for Grok 4.7.
  • GPT-6 Luna is cheaper at $0.10 / $0.50 per million input/output tokens, against $2 / $6 for Grok 4.7.
  • GPT-6 Luna accepts more context: 1.05M tokens versus 500K.

Side by side

GPT-6 Luna and Grok 4.7 specifications
GPT-6 LunaGrok 4.7
ProviderOpenAIxAI
Noometry Index53.353.1
Released2026-09-222026-09-21
WeightsProprietaryProprietary
Context window1.05M500K
Max output128K500K
Input $ / M tokens$0.10$2
Output $ / M tokens$0.50$6
Results tracked4239

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Grok 4.7 leads

GPT-6 Luna: 55.5 (#25), Grok 4.7: 58.0 (#18)

Coding benchmarks
BenchmarkGPT-6 LunaGrok 4.7
FrontierCode42.4%47.6%
LMArena WebDev15811639
SciCode54.6%57.8%
LMArena Coding14391427
DeepSWE66.6%—
CursorBench—46.3%
FrontierSWE—29.5%
ALE-Bench1,577—

Agentic & Tool Use Grok 4.7 leads

GPT-6 Luna: 33.3 (#54), Grok 4.7: 36.7 (#37)

Agentic & Tool Use benchmarks
BenchmarkGPT-6 LunaGrok 4.7
APEX-Agents44.3%54.6%
GDP.pdf23%22.8%
Vending-Bench 2—10,537

Reasoning Too close to call

GPT-6 Luna: 48.2 (#41), Grok 4.7: 49.1 (#40)

Reasoning benchmarks
BenchmarkGPT-6 LunaGrok 4.7
NYT Connections (extended)68.7%76.8%
CritPt19.4%18%
Chess Puzzles31%38%
LMArena Hard Prompts14111413
Mystery Game Puzzles7%29%
DTBench90.1%96%
LMCA44.5%49.4%
Epoch Capabilities Index156.28153.53
ARC-AGI-259.3%—
ARC-AGI-186.7%—

Math GPT-6 Luna leads

GPT-6 Luna: 76.1 (#15), Grok 4.7: 57.8 (#39)

Math benchmarks
BenchmarkGPT-6 LunaGrok 4.7
FrontierMath (Tiers 1-3)78.9%53%
FrontierMath Tier 456.1%17.1%
OTIS Mock AIME 2024-202598.9%98.1%
ProofBench64%34%
LMArena Math14161407

Knowledge Grok 4.7 leads

GPT-6 Luna: 57.0 (#41), Grok 4.7: 62.8 (#22)

Knowledge benchmarks
BenchmarkGPT-6 LunaGrok 4.7
GPQA Diamond90.5%92.7%
SimpleQA Verified41.4%56%
LMArena Expert14441422

Multimodal GPT-6 Luna leads

GPT-6 Luna: 42.4 (#30), Grok 4.7: 35.5 (#87)

Multimodal benchmarks
BenchmarkGPT-6 LunaGrok 4.7
LMArena Vision12171228
Blueprint-Bench 231.2%32.5%
Furniture Assembly44.2%20.8%

Multilingual Too close to call

GPT-6 Luna: 50.5 (#117), Grok 4.7: 50.8 (#116)

Multilingual benchmarks
BenchmarkGPT-6 LunaGrok 4.7
LMArena Non-English13861389
LMArena Chinese14331455
LMArena French14201455
LMArena Russian13941397
LMArena Spanish13931400
LMArena German1369—
LMArena Japanese1369—
LMArena Korean1360—

Instruction Following Too close to call

GPT-6 Luna: 74.3 (#99), Grok 4.7: 74.1 (#105)

Instruction Following benchmarks
BenchmarkGPT-6 LunaGrok 4.7
LMArena Instruction Following14091404

Long Context Too close to call

GPT-6 Luna: 43.0 (#111), Grok 4.7: 43.1 (#104)

Long Context benchmarks
BenchmarkGPT-6 LunaGrok 4.7
LMArena Longer Query14091413

Writing & Preference Grok 4.7 leads

GPT-6 Luna: 58.3 (#119), Grok 4.7: 70.0 (#24)

Writing & Preference benchmarks
BenchmarkGPT-6 LunaGrok 4.7
LMArena Text13911399
LMArena Creative Writing13631391
LMArena Multi-Turn13961393
EQ-Bench Creative Writing—2007

Frequently asked questions

Is GPT-6 Luna better than Grok 4.7?

GPT-6 Luna and Grok 4.7 score almost the same on the Noometry Index (53.3 vs 53.1), so choose on price, context window or the category you care about most.

Which is cheaper, GPT-6 Luna or Grok 4.7?

GPT-6 Luna is cheaper. It lists at $0.10 per million input tokens and $0.50 per million output tokens; Grok 4.7 lists at $2 and $6.

Is GPT-6 Luna or Grok 4.7 better for coding?

Grok 4.7 scores higher on coding benchmarks: 58.0 versus 55.5 in the Noometry coding category.

Which has the bigger context window?

GPT-6 Luna does, with 1.05M tokens against 500K.

How many benchmarks do GPT-6 Luna and Grok 4.7 share?

35 benchmarks have published results for both models. GPT-6 Luna has 42 scored results on Noometry and Grok 4.7 has 39.

Related comparisons

Go deeper