xAI, proprietary

Grok 4.5

Grok 4.5 by xAI ranks 25th of 354 ranked models on the Noometry Index as of October 2026, with a score of 55.0. Its strongest category is agentic & tool use, where it ranks 17th. API pricing starts at $2 per million input tokens and $6 per million output tokens, with a 500K-token context window.

Last verified

Specifications

Noometry rank
#25 of 354
Index score
55.0
Evidence
Confirmed 52 results
Provider
xAI
Released
July 8, 2026
Weights
Proprietary
Reasoning
Yes
Context window
500K
Max output
500K
Input price
$2 / M
Output price
$6 / M
Blended price
$3 / M
Output speed
4 tokens/s Kagi
Value
#155 of 219
Knowledge cutoff
Unknown
Input
text, image, pdf

Category scores

Each category score combines every public result we have in that category.

Grok 4.5 category scores
  1. Coding 52.2
  2. Agentic & Tool Use 44.4
  3. Reasoning 56.1
  4. Math 60.9
  5. Knowledge 62.3
  6. Multimodal 37.6
  7. Multilingual 54.4
  8. Instruction Following 76.0
  9. Long Context 44.8
  10. Writing & Preference 65.8
Grok 4.5 category ranks
CategoryScoreRankResults
Coding52.2#356
Agentic & Tool Use44.4#175
Reasoning56.1#2511
Math60.9#355
Knowledge62.3#243
Multimodal37.6#723
Multilingual54.4#421
Instruction Following76.0#481
Long Context44.8#561
Writing & Preference65.8#424

Strengths and weaknesses

Categories where Grok 4.5 places highest and lowest among the models ranked in each, with its score against that category's median.

Strongest categories

Grok 4.5: strongest categories
CategoryScorevs medianRank
Reasoning56.1+32.5#25 of 350, top 8%
Knowledge62.3+25.0#24 of 314, top 8%
Coding52.2+13.5#35 of 340, top 11%

Weakest categories

Grok 4.5: weakest categories
CategoryScorevs medianRank
Multimodal37.6−0.9#72 of 128, top 57%
Long Context44.8+3.8#56 of 296, top 19%
Instruction Following76.0+4.8#48 of 305, top 16%

Closest competitors

The models ranked just above and below Grok 4.5. When scores are this close, price and speed are often the better way to choose.

Models ranked closest to Grok 4.5
ModelRankScoreBlended $/MSpeed
Grok 4.6#2156.9$3—Compare
Qwen3.8 Max#2256.8$3—Compare
Gemini 3.1 Pro Preview#2356.7$4.50—Compare
Gemini 4 Argon#2456.5——Compare
GLM-5.3#2654.8$2.15—Compare
Muse Spark 1.3#2754.8$2—Compare
Gemini 3 Pro#2854.8—1Compare
Claude Sonnet 5#2954.6$4—Compare

Sponsored placements are available on pages like this one. Advertise on Noometry

Benchmark results

Every published result we track, with its source. Bold rows are the ones used for ranking; where several exist we prefer independent runs over self-reported numbers.

Coding

Grok 4.5 Coding benchmark results
BenchmarkScorePositionSettingSourceDate
DeepSWE53.8%#21 of 29, top 73%highEpoch AI
FrontierCode42.4%#17 of 37, top 46%Epoch AI
LMArena WebDev1553#33 of 113, top 30%LMArena2026-10-08
SciCode54.1%#29 of 121, top 24%highEpoch AI
WeirdML46.4%#56 of 119, top 48%Epoch AI
LMArena Coding1474#52 of 294, top 18%LMArena2026-10-08
ALE-Bench1,309#24 of 105, top 23%highEpoch AI

Agentic & Tool Use

Grok 4.5 Agentic & Tool Use benchmark results
BenchmarkScorePositionSettingSourceDate
APEX-Agents56.2%#17 of 49, top 35%Epoch AI
τ²-bench Banking47.9%#3 of 26, top 12%highτ²-bench2026-08-04
PostTrainBench23.4%#9 of 11, top 82%highEpoch AI
GBAEval65.4%#4 of 23, top 18%Epoch AI
GDP.pdf14%#31 of 36, top 87%highEpoch AI
LMArena Search1213#8 of 32, top 25%LMArena2026-08-24
Vending-Bench 23,887#36 of 60, top 60%Epoch AI

Reasoning

Grok 4.5 Reasoning benchmark results
BenchmarkScorePositionSettingSourceDate
ARC-AGI-252.6%#34 of 83, top 41%highEpoch AI
ARC-AGI-233.1%lowEpoch AI
ARC-AGI-252.6%mediumEpoch AI
SimpleBench70%#12 of 77, top 16%Epoch AI
Kagi LLM Benchmark83.5%#5 of 99, top 6%Kagi LLM Benchmark
NYT Connections (extended)79.9%#36 of 91, top 40%high reasoningLech Mazur benchmarks
ARC-AGI-185.7%highEpoch AI
ARC-AGI-179.2%lowEpoch AI
ARC-AGI-187.2%#32 of 83, top 39%mediumEpoch AI
CritPt15.4%#36 of 134, top 27%highEpoch AI
Chess Puzzles36%#28 of 129, top 22%highEpoch AI2026-07-08
LMArena Hard Prompts1462#43 of 297, top 15%LMArena2026-10-08
DTBench96.5%#10 of 151, top 7%highEpoch AI
LMCA45.2%#33 of 125, top 27%highEpoch AI
Surface Evolver Bench74.4%#8 of 25, top 32%highEpoch AI
Epoch Capabilities Index153.92#38 of 213, top 18%Epoch AI2026-07-08

Math

Grok 4.5 Math benchmark results
BenchmarkScorePositionSettingSourceDate
FrontierMath (Tiers 1-3)57.2%#39 of 81, top 49%highEpoch AI2026-07-09
FrontierMath Tier 424.4%#39 of 63, top 62%highEpoch AI2026-07-09
OTIS Mock AIME 2024-202597.8%#26 of 173, top 16%highEpoch AI2026-07-08
ProofBench31%#42 of 77, top 55%highEpoch AI
LMArena Math1459#52 of 285, top 19%LMArena2026-10-08

Knowledge

Grok 4.5 Knowledge benchmark results
BenchmarkScorePositionSettingSourceDate
GPQA Diamond93.4%#15 of 186, top 9%highEpoch AI2026-07-08
SimpleQA Verified48.3%#29 of 77, top 38%highEpoch AI2026-08-27
LMArena Expert1466#54 of 273, top 20%LMArena2026-10-08

Multimodal

Grok 4.5 Multimodal benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Vision1288#24 of 122, top 20%LMArena2026-10-09
Blueprint-Bench 227.3%#18 of 31, top 59%Epoch AI
Furniture Assembly22.5%#28 of 31, top 91%highEpoch AI2026-09-24
LMArena Document1452#21 of 38, top 56%LMArena2026-09-13

Multilingual

Grok 4.5 Multilingual benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Non-English1440#41 of 297, top 14%LMArena2026-10-08
LMArena Chinese1496#48 of 285, top 17%LMArena2026-10-08
LMArena French1456#57 of 223, top 26%LMArena2026-10-08
LMArena German1446#49 of 231, top 22%LMArena2026-10-08
LMArena Japanese1428#34 of 211, top 17%LMArena2026-10-08
LMArena Korean1404#49 of 213, top 24%LMArena2026-10-08
LMArena Russian1448#46 of 283, top 17%LMArena2026-10-08
LMArena Spanish1450#50 of 226, top 23%LMArena2026-10-08

Instruction Following

Grok 4.5 Instruction Following benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Instruction Following1446#44 of 298, top 15%LMArena2026-10-08

Long Context

Grok 4.5 Long Context benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Longer Query1463#39 of 291, top 14%LMArena2026-10-08

Writing & Preference

Grok 4.5 Writing & Preference benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Text1448#47 of 297, top 16%LMArena2026-10-08
LMArena Creative Writing1442#34 of 295, top 12%LMArena2026-10-08
EQ-Bench Creative Writing1579#45 of 115, top 40%EQ-Bench
LMArena Multi-Turn1456#44 of 295, top 15%LMArena2026-10-08

API pricing by provider

Grok 4.5 API prices
RouteInput $/MOutput $/MCached input $/MChecked
openrouter$2$6$0.302026-10-10
xai$2$6$0.302026-10-10

Compare Grok 4.5

Other xAI models

Frequently asked questions

How good is Grok 4.5?

Grok 4.5 by xAI ranks 25th of 354 ranked models on the Noometry Index as of October 2026, with a score of 55.0. Its strongest category is agentic & tool use, where it ranks 17th. API pricing starts at $2 per million input tokens and $6 per million output tokens, with a 500K-token context window.

How much does Grok 4.5 cost?

Grok 4.5 costs $2 per million input tokens and $6 per million output tokens on xAI's own API, with cached input at $0.30.

What is Grok 4.5's context window?

Grok 4.5 accepts up to 500K tokens of input and can write up to 500K tokens in one response.

Is Grok 4.5 open source?

No. Grok 4.5 is proprietary and available only through xAI's API and partner platforms.

How fast is Grok 4.5?

Grok 4.5 generated about 4 output tokens per second in the Kagi LLM Benchmark's timed runs. Speed varies by provider, load and reasoning effort.

What are Grok 4.5's strengths and weaknesses?

Relative to other ranked models, Grok 4.5 places best in reasoning, knowledge, coding and lowest in multimodal, long context, instruction following.

What is Grok 4.5 best at?

Its best category is agentic & tool use, where it ranks 17th on Noometry.