Model comparison

GPT-5.2 vs Grok 4.6

Grok 4.6 is the stronger model overall, scoring 56.9 to 54.1 on the Noometry Index.

Last verified . 40 shared benchmarks.

GPT-5.2 OpenAI

54.1

Rank #34 Confirmed

Grok 4.6 xAI

56.9

Rank #21 Confirmed

Summary

  • They share 40 benchmarks with published results for both. GPT-5.2 scores higher in 4 categories and Grok 4.6 in 6 categories; 6 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Grok 4.6 leads 61.4 to 50.2.
  • The biggest single-benchmark swing is ProofBench: 15% for GPT-5.2 and 51% for Grok 4.6.
  • Grok 4.6 is cheaper at $2 / $6 per million input/output tokens, against $1.75 / $14 for GPT-5.2.
  • Grok 4.6 accepts more context: 500K tokens versus 400K.

Side by side

GPT-5.2 and Grok 4.6 specifications
GPT-5.2Grok 4.6
ProviderOpenAIxAI
Noometry Index54.156.9
Released2025-12-112026-08-12
WeightsProprietaryProprietary
Context window400K500K
Max output128K500K
Input $ / M tokens$1.75$2
Output $ / M tokens$14$6
Results tracked6749

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Grok 4.6 leads

GPT-5.2: 51.6 (#37), Grok 4.6: 58.5 (#16)

Coding benchmarks
BenchmarkGPT-5.2Grok 4.6
LMArena WebDev14161617
WeirdML72.2%67.3%
LMArena Coding14471465
ALE-Bench1,2941,508
SWE-bench Verified73.8%—
DeepSWE—67.5%
FrontierCode—48%
SWE-bench Verified (bash only)72.8%—
CursorBench—41.4%
SWE-bench Multilingual66.7%—
FrontierSWE—25.3%
SciCode—56.5%
GSO27.4%—
AlgoTune2.05—

Agentic & Tool Use Too close to call

GPT-5.2: 40.2 (#24), Grok 4.6: 39.4 (#27)

Agentic & Tool Use benchmarks
BenchmarkGPT-5.2Grok 4.6
Vending-Bench 23,5919,047
Terminal-Bench64.9%—
APEX-Agents—65.3%
Berkeley Function Calling Leaderboard55.9%—
GDPval49.7%—
Remote Labor Index2.5%—
τ²-bench Airline83%—
τ²-bench Banking32.2%—
τ²-bench Retail81.6%—
τ²-bench Telecom89.7%—
DeepResearch Bench41.1%—
GDP.pdf—17.2%
LMArena Search1207—
METR Time Horizons75.3%—

Reasoning Grok 4.6 leads

GPT-5.2: 50.2 (#35), Grok 4.6: 61.4 (#20)

Reasoning benchmarks
BenchmarkGPT-5.2Grok 4.6
ARC-AGI-252.9%67.1%
SimpleBench45.8%75.9%
NYT Connections (extended)83.6%80%
ARC-AGI-186.2%87.5%
Chess Puzzles49%40%
EBR-Bench23%30.5%
LMArena Hard Prompts14451447
Mystery Game Puzzles23%34%
DTBench90.9%97.3%
LMCA43.9%48.5%
Epoch Capabilities Index153.45156.44
Kagi LLM Benchmark73.3%—
CritPt—19.7%
EnigmaEval10.4%—
ForecastBench60.1—

Math Grok 4.6 leads

GPT-5.2: 60.0 (#38), Grok 4.6: 67.0 (#24)

Knowledge Grok 4.6 leads

GPT-5.2: 59.3 (#32), Grok 4.6: 63.3 (#20)

Knowledge benchmarks
BenchmarkGPT-5.2Grok 4.6
GPQA Diamond91.4%94%
SimpleQA Verified37.1%49.3%
LMArena Expert14451467
Humanity's Last Exam27.8%—
Vectara Hallucination Rate8.4%—

Multimodal GPT-5.2 leads

GPT-5.2: 51.3 (#7), Grok 4.6: 43.6 (#23)

Multimodal benchmarks
BenchmarkGPT-5.2Grok 4.6
LMArena Vision12681263
Furniture Assembly38.3%40%
LMArena Document14051452
VPCT84%—
Blueprint-Bench 2—33.2%

Multilingual Too close to call

GPT-5.2: 53.4 (#67), Grok 4.6: 53.0 (#74)

Multilingual benchmarks
BenchmarkGPT-5.2Grok 4.6
LMArena Non-English14251420
LMArena Chinese14601480
LMArena French14551461
LMArena German14481431
LMArena Japanese14201376
LMArena Korean13921397
LMArena Russian14401422
LMArena Spanish14331404

Instruction Following Too close to call

GPT-5.2: 74.7 (#89), Grok 4.6: 75.4 (#63)

Instruction Following benchmarks
BenchmarkGPT-5.2Grok 4.6
LMArena Instruction Following14171431

Long Context Too close to call

GPT-5.2: 44.0 (#78), Grok 4.6: 44.5 (#66)

Long Context benchmarks
BenchmarkGPT-5.2Grok 4.6
LMArena Longer Query14281454
CL-bench18.2%—

Writing & Preference GPT-5.2 leads

GPT-5.2: 66.8 (#32), Grok 4.6: 62.3 (#80)

Writing & Preference benchmarks
BenchmarkGPT-5.2Grok 4.6
LMArena Text14391428
LMArena Creative Writing14011428
LMArena Multi-Turn14581425
EQ-Bench Creative Writing1703—

Frequently asked questions

Is GPT-5.2 better than Grok 4.6?

Grok 4.6 is the stronger model overall, scoring 56.9 to 54.1 on the Noometry Index.

Which is cheaper, GPT-5.2 or Grok 4.6?

Grok 4.6 is cheaper. It lists at $2 per million input tokens and $6 per million output tokens; GPT-5.2 lists at $1.75 and $14.

Is GPT-5.2 or Grok 4.6 better for coding?

Grok 4.6 scores higher on coding benchmarks: 58.5 versus 51.6 in the Noometry coding category.

Which has the bigger context window?

Grok 4.6 does, with 500K tokens against 400K.

How many benchmarks do GPT-5.2 and Grok 4.6 share?

40 benchmarks have published results for both models. GPT-5.2 has 67 scored results on Noometry and Grok 4.6 has 49.

Related comparisons

Go deeper