xAI, proprietary

Grok 4.6

Grok 4.6 by xAI ranks 21st of 354 ranked models on the Noometry Index as of October 2026, with a score of 56.9. Its strongest category is coding, where it ranks 16th. API pricing starts at $2 per million input tokens and $6 per million output tokens, with a 500K-token context window.

Last verified

Specifications

Noometry rank
#21 of 354
Index score
56.9
Evidence
Confirmed 49 results
Provider
xAI
Released
August 12, 2026
Weights
Proprietary
Reasoning
Yes
Context window
500K
Max output
500K
Input price
$2 / M
Output price
$6 / M
Blended price
$3 / M
Output speed
Not measured
Value
#153 of 219
Knowledge cutoff
February 2026
Input
text, image, pdf

Category scores

Each category score combines every public result we have in that category.

Grok 4.6 category scores
  1. Coding 58.5
  2. Agentic & Tool Use 39.4
  3. Reasoning 61.4
  4. Math 67.0
  5. Knowledge 63.3
  6. Multimodal 43.6
  7. Multilingual 53.0
  8. Instruction Following 75.4
  9. Long Context 44.5
  10. Writing & Preference 62.3
Grok 4.6 category ranks
CategoryScoreRankResults
Coding58.5#168
Agentic & Tool Use39.4#272
Reasoning61.4#2011
Math67.0#245
Knowledge63.3#203
Multimodal43.6#233
Multilingual53.0#741
Instruction Following75.4#631
Long Context44.5#661
Writing & Preference62.3#803

Strengths and weaknesses

Categories where Grok 4.6 places highest and lowest among the models ranked in each, with its score against that category's median.

Strongest categories

Grok 4.6: strongest categories
CategoryScorevs medianRank
Coding58.5+19.8#16 of 340, top 5%
Reasoning61.4+37.8#20 of 350, top 6%
Knowledge63.3+26.0#20 of 314, top 7%

Weakest categories

Grok 4.6: weakest categories
CategoryScorevs medianRank
Writing & Preference62.3+8.5#80 of 312, top 26%
Multilingual53.0+5.6#74 of 297, top 25%
Long Context44.5+3.6#66 of 296, top 23%

Closest competitors

The models ranked just above and below Grok 4.6. When scores are this close, price and speed are often the better way to choose.

Models ranked closest to Grok 4.6
ModelRankScoreBlended $/MSpeed
GPT-5.6 Terra#1759.2$4.5011Compare
GPT-5.4 Pro#1858.9$67.50—Compare
Claude Opus 4.7#1958.3$1033Compare
Claude Opus 4.6#2058.2$1019Compare
Qwen3.8 Max#2256.8$3—Compare
Gemini 3.1 Pro Preview#2356.7$4.50—Compare
Gemini 4 Argon#2456.5——Compare
Grok 4.5#2555.0$34Compare

Sponsored placements are available on pages like this one. Advertise on Noometry

Benchmark results

Every published result we track, with its source. Bold rows are the ones used for ranking; where several exist we prefer independent runs over self-reported numbers.

Coding

Grok 4.6 Coding benchmark results
BenchmarkScorePositionSettingSourceDate
DeepSWE65.2%highEpoch AI
DeepSWE41.6%lowEpoch AI
DeepSWE67.5%#11 of 29, top 38%mediumEpoch AI
DeepSWE66.7%xhighEpoch AI
FrontierCode48%#9 of 37, top 25%Epoch AI
CursorBench40.4%highEpoch AI
CursorBench33.4%lowEpoch AI
CursorBench36.1%mediumEpoch AI
CursorBench41.4%#9 of 14, top 65%xhighEpoch AI
LMArena WebDev1617#20 of 113, top 18%highLMArena2026-10-08
FrontierSWE25.3%#12 of 18, top 67%xhighEpoch AI
SciCode56.5%#20 of 121, top 17%highEpoch AI
SciCode48.4%lowEpoch AI
SciCode54.6%mediumEpoch AI
SciCode51.6%xhighEpoch AI
WeirdML67.3%#21 of 119, top 18%highEpoch AI
LMArena Coding1465#66 of 294, top 23%highLMArena2026-10-08
ALE-Bench1,508#17 of 105, top 17%xhighEpoch AI

Agentic & Tool Use

Grok 4.6 Agentic & Tool Use benchmark results
BenchmarkScorePositionSettingSourceDate
APEX-Agents65.3%#7 of 49, top 15%Epoch AI
GDP.pdf16%highEpoch AI
GDP.pdf17.2%#23 of 36, top 64%xhighEpoch AI
Vending-Bench 29,047#9 of 60, top 15%Epoch AI

Reasoning

Grok 4.6 Reasoning benchmark results
BenchmarkScorePositionSettingSourceDate
ARC-AGI-265.1%highEpoch AI
ARC-AGI-227.6%lowEpoch AI
ARC-AGI-261.3%mediumEpoch AI
ARC-AGI-267.1%#22 of 83, top 27%xhighEpoch AI
SimpleBench75.9%#7 of 77, top 10%Epoch AI
NYT Connections (extended)80%#35 of 91, top 39%xhigh reasoningLech Mazur benchmarks
ARC-AGI-187%highEpoch AI
ARC-AGI-174.8%lowEpoch AI
ARC-AGI-187.5%#30 of 83, top 37%mediumEpoch AI
ARC-AGI-187%xhighEpoch AI
CritPt17.1%highEpoch AI
CritPt5.7%lowEpoch AI
CritPt17.7%mediumEpoch AI
CritPt19.7%#25 of 134, top 19%xhighEpoch AI
Chess Puzzles40%#21 of 129, top 17%highEpoch AI2026-08-12
Chess Puzzles31%xhighEpoch AI2026-08-14
EBR-Bench30.5%#10 of 24, top 42%xhighEpoch AI2026-08-18
LMArena Hard Prompts1447#68 of 297, top 23%highLMArena2026-10-08
Mystery Game Puzzles34%#22 of 74, top 30%xhighEpoch AI2026-08-14
DTBench97.3%#7 of 151, top 5%xhighEpoch AI
LMCA48.5%#25 of 125, top 20%xhighEpoch AI
Epoch Capabilities Index156.44#21 of 213, top 10%Epoch AI2026-08-12

Math

Grok 4.6 Math benchmark results
BenchmarkScorePositionSettingSourceDate
FrontierMath (Tiers 1-3)66%#30 of 81, top 38%xhighEpoch AI2026-08-14
FrontierMath Tier 431.7%#27 of 63, top 43%xhighEpoch AI2026-08-14
OTIS Mock AIME 2024-202597.8%highEpoch AI2026-08-12
OTIS Mock AIME 2024-202599.2%#13 of 173, top 8%xhighEpoch AI2026-08-14
ProofBench51%#27 of 77, top 36%Epoch AI
LMArena Math1423#96 of 285, top 34%highLMArena2026-10-08

Knowledge

Grok 4.6 Knowledge benchmark results
BenchmarkScorePositionSettingSourceDate
GPQA Diamond94%#11 of 186, top 6%highEpoch AI2026-08-12
GPQA Diamond93.2%xhighEpoch AI2026-08-14
SimpleQA Verified49.3%#27 of 77, top 36%highEpoch AI2026-08-27
SimpleQA Verified48.9%xhighEpoch AI2026-08-27
LMArena Expert1467#52 of 273, top 20%highLMArena2026-10-08

Multimodal

Grok 4.6 Multimodal benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Vision1263#44 of 122, top 37%highLMArena2026-10-09
Blueprint-Bench 233.2%#11 of 31, top 36%Epoch AI
Furniture Assembly40%#15 of 31, top 49%xhighEpoch AI2026-09-24
LMArena Document1452#20 of 38, top 53%highLMArena2026-09-13

Multilingual

Grok 4.6 Multilingual benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Non-English1420#74 of 297, top 25%highLMArena2026-10-08
LMArena Chinese1480#64 of 285, top 23%highLMArena2026-10-08
LMArena French1461#47 of 223, top 22%highLMArena2026-10-08
LMArena German1431#64 of 231, top 28%highLMArena2026-10-08
LMArena Japanese1376#83 of 211, top 40%highLMArena2026-10-08
LMArena Korean1397#57 of 213, top 27%highLMArena2026-10-08
LMArena Russian1422#79 of 283, top 28%highLMArena2026-10-08
LMArena Spanish1404#106 of 226, top 47%highLMArena2026-10-08

Instruction Following

Grok 4.6 Instruction Following benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Instruction Following1431#59 of 298, top 20%highLMArena2026-10-08

Long Context

Grok 4.6 Long Context benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Longer Query1454#49 of 291, top 17%highLMArena2026-10-08

Writing & Preference

Grok 4.6 Writing & Preference benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Text1428#80 of 297, top 27%highLMArena2026-10-08
LMArena Creative Writing1428#49 of 295, top 17%highLMArena2026-10-08
LMArena Multi-Turn1425#90 of 295, top 31%highLMArena2026-10-08

API pricing by provider

Grok 4.6 API prices
RouteInput $/MOutput $/MCached input $/MChecked
azure$2$6$0.502026-10-10
bedrock$2$6$0.502026-10-10
openrouter$2$6$0.502026-10-10
vertex$2$6$0.502026-10-10
xai$2$6$0.502026-10-10

Compare Grok 4.6

Other xAI models

Frequently asked questions

How good is Grok 4.6?

Grok 4.6 by xAI ranks 21st of 354 ranked models on the Noometry Index as of October 2026, with a score of 56.9. Its strongest category is coding, where it ranks 16th. API pricing starts at $2 per million input tokens and $6 per million output tokens, with a 500K-token context window.

How much does Grok 4.6 cost?

Grok 4.6 costs $2 per million input tokens and $6 per million output tokens on xAI's own API, with cached input at $0.50.

What is Grok 4.6's context window?

Grok 4.6 accepts up to 500K tokens of input and can write up to 500K tokens in one response.

Is Grok 4.6 open source?

No. Grok 4.6 is proprietary and available only through xAI's API and partner platforms.

What are Grok 4.6's strengths and weaknesses?

Relative to other ranked models, Grok 4.6 places best in coding, reasoning, knowledge and lowest in writing & preference, multilingual, long context.

What is Grok 4.6 best at?

Its best category is coding, where it ranks 16th on Noometry.