xAI, proprietary
Grok 4
Grok 4 by xAI ranks 56th of 354 ranked models on the Noometry Index as of October 2026, with a score of 48.1. Its strongest category is long context, where it ranks 4th.
Last verified
Specifications
- Noometry rank
- #56 of 354
- Index score
- 48.1
- Evidence
- Confirmed 48 results
- Provider
- xAI
- Released
- July 9, 2025
- Weights
- Proprietary
- Reasoning
- Unknown
- Context window
- —
- Max output
- —
- Input price
- Not listed
- Output price
- Not listed
- Blended price
- Not listed
- Output speed
- 1 tokens/s Kagi
- Value
- Not ranked
- Knowledge cutoff
- Unknown
Category scores
Each category score combines every public result we have in that category.
- Coding 50.3
- Agentic & Tool Use 32.3
- Reasoning 36.7
- Math 48.4
- Knowledge 53.8
- Multimodal 33.7
- Multilingual 51.8
- Instruction Following 79.2
- Long Context 63.1
- Writing & Preference 58.5
| Category | Score | Rank | Results |
|---|---|---|---|
| Coding | 50.3 | #46 | 3 |
| Agentic & Tool Use | 32.3 | #68 | 6 |
| Reasoning | 36.7 | #65 | 6 |
| Math | 48.4 | #64 | 3 |
| Knowledge | 53.8 | #55 | 5 |
| Multimodal | 33.7 | #94 | 2 |
| Multilingual | 51.8 | #103 | 1 |
| Instruction Following | 79.2 | #5 | 2 |
| Long Context | 63.1 | #4 | 2 |
| Writing & Preference | 58.5 | #116 | 5 |
Strengths and weaknesses
Categories where Grok 4 places highest and lowest among the models ranked in each, with its score against that category's median.
Strongest categories
| Category | Score | vs median | Rank |
|---|---|---|---|
| Long Context | 63.1 | +22.2 | #4 of 296, top 2% |
| Instruction Following | 79.2 | +8.0 | #5 of 305, top 2% |
| Coding | 50.3 | +11.6 | #46 of 340, top 14% |
Weakest categories
| Category | Score | vs median | Rank |
|---|---|---|---|
| Multimodal | 33.7 | −4.8 | #94 of 128, top 74% |
| Agentic & Tool Use | 32.3 | +1.9 | #68 of 154, top 45% |
| Writing & Preference | 58.5 | +4.7 | #116 of 312, top 38% |
Closest competitors
The models ranked just above and below Grok 4. When scores are this close, price and speed are often the better way to choose.
| Model | Rank | Score | Blended $/M | Speed | |
|---|---|---|---|---|---|
| Claude Haiku 5.5 | #52 | 49.5 | $0.20 | — | Compare |
| GPT-5.1 | #53 | 49.0 | $3.44 | — | Compare |
| Grok 4.20 (Non-Reasoning) | #54 | 48.6 | $1.56 | 61 | Compare |
| MiMo-V2.6-Flash | #55 | 48.5 | $0.18 | — | Compare |
| Kimi K2.5 | #57 | 48.1 | $0.90 | 66 | Compare |
| Step 5 Preview | #58 | 47.9 | $1.43 | — | Compare |
| GLM-5.1 | #59 | 47.8 | $2.15 | — | Compare |
| Kimi K2.6 | #60 | 47.7 | $1.71 | — | Compare |
Sponsored placements are available on pages like this one. Advertise on Noometry
Benchmark results
Every published result we track, with its source. Bold rows are the ones used for ranking; where several exist we prefer independent runs over self-reported numbers.
Coding
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| Aider Polyglot | 79.6% | #5 of 44, top 12% | Epoch AI | ||
| Aider Polyglot | 79.6% | high | Epoch AI | ||
| WeirdML | 45.7% | #59 of 119, top 50% | Epoch AI | ||
| LMArena Coding | 1408 | #130 of 294, top 45% | LMArena | 2026-10-08 |
Agentic & Tool Use
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| Terminal-Bench | 27.2% | #33 of 41, top 81% | Epoch AI | ||
| Berkeley Function Calling Leaderboard | 63% | #8 of 49, top 17% | prompt | Berkeley Function Calling Leaderboard | |
| GDPval | 21.1% | #10 of 11, top 91% | high | Epoch AI | |
| Cybench | 43% | #4 of 21, top 20% | Epoch AI | ||
| DeepResearch Bench | 47.3% | #11 of 24, top 46% | Epoch AI | ||
| BALROG | 43.6% | #9 of 35, top 26% | Epoch AI | ||
| LMArena Search | 1142 | #27 of 32, top 85% | LMArena | 2026-08-24 | |
| METR Time Horizons | 66.6% | #13 of 32, top 41% | Epoch AI |
Reasoning
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| ARC-AGI-2 | 16% | #45 of 83, top 55% | Epoch AI | ||
| SimpleBench | 60.5% | #25 of 77, top 33% | Epoch AI | ||
| Kagi LLM Benchmark | 73.6% | #16 of 99, top 17% | Kagi LLM Benchmark | ||
| ARC-AGI-1 | 66.7% | #44 of 83, top 54% | Epoch AI | ||
| Chess Puzzles | 28% | #38 of 129, top 30% | Epoch AI | 2026-01-30 | |
| LMArena Hard Prompts | 1409 | #119 of 297, top 41% | LMArena | 2026-10-08 | |
| Epoch Capabilities Index | 146.44 | #73 of 213, top 35% | Epoch AI | 2025-07-09 | |
| ForecastBench | 60.9 | #21 of 72, top 30% | Epoch AI |
Math
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| OTIS Mock AIME 2024-2025 | 84% | #74 of 173, top 43% | Epoch AI | ||
| Omni-MATH | 60.3% | #9 of 57, top 16% | HELM Capabilities | ||
| LMArena Math | 1422 | #98 of 285, top 35% | LMArena | 2026-10-08 | |
| FrontierMath (Feb 2025 set) | 19.7% | #31 of 68, top 46% | Epoch AI | 2025-11-13 | |
| FrontierMath Tier 4 (v1) | 2.1% | #41 of 55, top 75% | Epoch AI | 2025-08-11 |
Knowledge
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| GPQA Diamond | 87% | #56 of 186, top 31% | Epoch AI | ||
| MMLU-Pro | 85.1% | #7 of 58, top 13% | HELM Capabilities | ||
| Confabulations (lower is better) | 12.4% | #7 of 51, top 14% | Lech Mazur benchmarks | ||
| GPQA (HELM) | 72.7% | #7 of 57, top 13% | HELM Capabilities | ||
| LMArena Expert | 1415 | #111 of 273, top 41% | LMArena | 2026-10-08 |
Multimodal
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Vision | 1210 | #76 of 122, top 63% | LMArena | 2026-10-09 | |
| GeoBench | 45% | #22 of 25, top 88% | Epoch AI |
Multilingual
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Non-English | 1403 | #103 of 297, top 35% | LMArena | 2026-10-08 | |
| LMArena Chinese | 1427 | #120 of 285, top 43% | LMArena | 2026-10-08 | |
| LMArena French | 1418 | #103 of 223, top 47% | LMArena | 2026-10-08 | |
| LMArena German | 1429 | #67 of 231, top 30% | LMArena | 2026-10-08 | |
| LMArena Japanese | 1394 | #65 of 211, top 31% | LMArena | 2026-10-08 | |
| LMArena Korean | 1377 | #82 of 213, top 39% | LMArena | 2026-10-08 | |
| LMArena Russian | 1410 | #95 of 283, top 34% | LMArena | 2026-10-08 | |
| LMArena Spanish | 1420 | #92 of 226, top 41% | LMArena | 2026-10-08 |
Instruction Following
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| IFEval | 94.9% | #2 of 57, top 4% | HELM Capabilities | ||
| LMArena Instruction Following | 1387 | #116 of 298, top 39% | LMArena | 2026-10-08 |
Long Context
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| Fiction.LiveBench | 94.4% | #3 of 47, top 7% | Epoch AI | ||
| LMArena Longer Query | 1409 | #109 of 291, top 38% | LMArena | 2026-10-08 |
Writing & Preference
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Text | 1411 | #108 of 297, top 37% | LMArena | 2026-10-08 | |
| LMArena Creative Writing | 1397 | #85 of 295, top 29% | LMArena | 2026-10-08 | |
| Short-Story Creative Writing | 76.9% | #20 of 39, top 52% | Epoch AI | ||
| WildBench | 79.7% | #32 of 57, top 57% | HELM Capabilities | ||
| LMArena Multi-Turn | 1416 | #99 of 295, top 34% | LMArena | 2026-10-08 |
Compare Grok 4
- Grok 4 vs Grok 3
- Grok 4 vs MiMo-V2.6-Flash
- Grok 4 vs Kimi K2.5
- Grok 4 vs Grok 4.20 (Non-Reasoning)
- Grok 4 vs Step 5 Preview
- Grok 4 vs GPT-5.1
- Grok 4 vs GLM-5.1
- Grok 4 vs GPT-6 Astra
- Grok 4 vs Claude Fable 5.1
- Grok 4 vs Gemini 3.8 Flash
- Grok 4 vs Kimi K3
- Grok 4 vs Qwen3.8 Max
- Grok 4 vs GLM-5.3
- Grok 4 vs Muse Spark 1.3
Other xAI models
- Grok 4.656.9
- Grok 4.555.0
- Grok 4.753.1
- Grok 4.20 (Non-Reasoning)48.6
- Grok 4.20 Multi-Agent46.2
- Grok 4.343.8
- Grok 4.141.5
- Grok 4.1 Fast41.4
Frequently asked questions
How good is Grok 4?
Grok 4 by xAI ranks 56th of 354 ranked models on the Noometry Index as of October 2026, with a score of 48.1. Its strongest category is long context, where it ranks 4th.
Is Grok 4 open source?
No. Grok 4 is proprietary and available only through xAI's API and partner platforms.
How fast is Grok 4?
Grok 4 generated about 1 output tokens per second in the Kagi LLM Benchmark's timed runs. Speed varies by provider, load and reasoning effort.
What are Grok 4's strengths and weaknesses?
Relative to other ranked models, Grok 4 places best in long context, instruction following, coding and lowest in multimodal, agentic & tool use, writing & preference.
What is Grok 4 best at?
Its best category is long context, where it ranks 4th on Noometry.