xAI, proprietary
Grok 3
Grok 3 by xAI ranks 157th of 354 ranked models on the Noometry Index as of October 2026, with a score of 39.9. Its strongest category is instruction following, where it ranks 73rd.
Last verified
Specifications
- Noometry rank
- #157 of 354
- Index score
- 39.9
- Evidence
- Confirmed 40 results
- Provider
- xAI
- Released
- April 9, 2025
- Weights
- Proprietary
- Reasoning
- Unknown
- Context window
- —
- Max output
- —
- Input price
- Not listed
- Output price
- Not listed
- Blended price
- Not listed
- Output speed
- 42 tokens/s Kagi
- Value
- Not ranked
- Knowledge cutoff
- Unknown
Category scores
Each category score combines every public result we have in that category.
- Coding 41.9
- Agentic & Tool Use 30.5
- Reasoning 13.7
- Math 38.0
- Knowledge 46.2
- Multilingual 52.3
- Instruction Following 75.0
- Long Context 38.7
- Writing & Preference 55.8
| Category | Score | Rank | Results |
|---|---|---|---|
| Coding | 41.9 | #115 | 3 |
| Agentic & Tool Use | 30.5 | #76 | 1 |
| Reasoning | 13.7 | #333 | 5 |
| Math | 38.0 | #145 | 4 |
| Knowledge | 46.2 | #82 | 6 |
| Multilingual | 52.3 | #87 | 1 |
| Instruction Following | 75.0 | #73 | 2 |
| Long Context | 38.7 | #192 | 2 |
| Writing & Preference | 55.8 | #141 | 6 |
Strengths and weaknesses
Categories where Grok 3 places highest and lowest among the models ranked in each, with its score against that category's median.
Strongest categories
| Category | Score | vs median | Rank |
|---|---|---|---|
| Instruction Following | 75.0 | +3.8 | #73 of 305, top 24% |
| Knowledge | 46.2 | +8.8 | #82 of 314, top 27% |
| Multilingual | 52.3 | +4.9 | #87 of 297, top 30% |
Weakest categories
| Category | Score | vs median | Rank |
|---|---|---|---|
| Reasoning | 13.7 | −10.0 | #333 of 350, top 96% |
| Long Context | 38.7 | −2.2 | #192 of 296, top 65% |
| Agentic & Tool Use | 30.5 | +0.2 | #76 of 154, top 50% |
Closest competitors
The models ranked just above and below Grok 3. When scores are this close, price and speed are often the better way to choose.
| Model | Rank | Score | Blended $/M | Speed | |
|---|---|---|---|---|---|
| Nemotron 3 Super | #153 | 40.1 | $0.17 | — | Compare |
| Llama 3.3 Nemotron 49b Super v1 | #154 | 40.1 | — | — | Compare |
| Nemotron 3.5 Lightning | #155 | 40.0 | $0.0875 | — | Compare |
| Qwen3.7 Flash | #156 | 39.9 | $0.055 | — | Compare |
| GLM-4.5V | #158 | 39.8 | $0.90 | 34 | Compare |
| QwQ-32B | #159 | 39.8 | — | — | Compare |
| Step 1o Turbo 202506 | #160 | 39.7 | — | — | Compare |
| Nova 2 Lite | #161 | 39.7 | $0.85 | — | Compare |
Sponsored placements are available on pages like this one. Advertise on Noometry
Benchmark results
Every published result we track, with its source. Bold rows are the ones used for ranking; where several exist we prefer independent runs over self-reported numbers.
Coding
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| Aider Polyglot | 53.3% | #20 of 44, top 46% | Epoch AI | ||
| WeirdML | 37.2% | #87 of 119, top 74% | Epoch AI | ||
| LMArena Coding | 1432 | #109 of 294, top 38% | LMArena | 2026-10-08 |
Agentic & Tool Use
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| BALROG | 29.5% | #18 of 35, top 52% | Epoch AI |
Reasoning
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| ARC-AGI-2 | 0% | #79 of 83, top 96% | Epoch AI | ||
| SimpleBench | 36.1% | #56 of 77, top 73% | Epoch AI | ||
| Kagi LLM Benchmark | 61.3% | #40 of 99, top 41% | Kagi LLM Benchmark | ||
| ARC-AGI-1 | 5.5% | #77 of 83, top 93% | Epoch AI | ||
| LMArena Hard Prompts | 1434 | #87 of 297, top 30% | LMArena | 2026-10-08 | |
| Epoch Capabilities Index | 138.33 | #114 of 213, top 54% | Epoch AI | 2025-04-09 |
Math
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| OTIS Mock AIME 2024-2025 | 55.6% | #110 of 173, top 64% | Epoch AI | 2025-04-10 | |
| Omni-MATH | 46.4% | #20 of 57, top 36% | HELM Capabilities | ||
| LMArena Math | 1391 | #135 of 285, top 48% | LMArena | 2026-10-08 | |
| MATH Level 5 | 88.7% | #17 of 79, top 22% | Epoch AI | 2025-04-10 | |
| FrontierMath (Feb 2025 set) | 3.8% | #52 of 68, top 77% | Epoch AI | 2025-04-10 | |
| FrontierMath Tier 4 (v1) | 0% | #51 of 55, top 93% | Epoch AI | 2025-07-01 |
Knowledge
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| GPQA Diamond | 75.8% | #93 of 186, top 50% | Epoch AI | 2025-05-26 | |
| MMLU-Pro | 78.8% | #19 of 58, top 33% | HELM Capabilities | ||
| Confabulations (lower is better) | 14.2% | #14 of 51, top 28% | no reasoning | Lech Mazur benchmarks | |
| Vectara Hallucination Rate (lower is better) | 5.8% | #20 of 96, top 21% | Vectara Hallucination Leaderboard | ||
| GPQA (HELM) | 65% | #18 of 57, top 32% | HELM Capabilities | ||
| LMArena Expert | 1421 | #106 of 273, top 39% | LMArena | 2026-10-08 |
Multilingual
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Non-English | 1410 | #87 of 297, top 30% | LMArena | 2026-10-08 | |
| LMArena Chinese | 1448 | #103 of 285, top 37% | LMArena | 2026-10-08 | |
| LMArena French | 1460 | #49 of 223, top 22% | LMArena | 2026-10-08 | |
| LMArena German | 1431 | #65 of 231, top 29% | LMArena | 2026-10-08 | |
| LMArena Japanese | 1387 | #71 of 211, top 34% | LMArena | 2026-10-08 | |
| LMArena Korean | 1373 | #83 of 213, top 39% | LMArena | 2026-10-08 | |
| LMArena Russian | 1416 | #86 of 283, top 31% | LMArena | 2026-10-08 | |
| LMArena Spanish | 1417 | #94 of 226, top 42% | LMArena | 2026-10-08 |
Instruction Following
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| IFEval | 88.4% | #12 of 57, top 22% | HELM Capabilities | ||
| LMArena Instruction Following | 1409 | #89 of 298, top 30% | LMArena | 2026-10-08 |
Long Context
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| Fiction.LiveBench | 58.3% | #31 of 47, top 66% | Epoch AI | ||
| LMArena Longer Query | 1439 | #66 of 291, top 23% | LMArena | 2026-10-08 |
Writing & Preference
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Text | 1426 | #86 of 297, top 29% | LMArena | 2026-10-08 | |
| LMArena Creative Writing | 1414 | #60 of 295, top 21% | LMArena | 2026-10-08 | |
| Short-Story Creative Writing | 76.4% | #22 of 39, top 57% | Epoch AI | ||
| EQ-Bench Creative Writing | 1186 | #86 of 115, top 75% | EQ-Bench | ||
| WildBench | 84.9% | #13 of 57, top 23% | HELM Capabilities | ||
| LMArena Multi-Turn | 1425 | #91 of 295, top 31% | LMArena | 2026-10-08 |
Compare Grok 3
- Grok 3 vs Qwen3.7 Flash
- Grok 3 vs GLM-4.5V
- Grok 3 vs Nemotron 3.5 Lightning
- Grok 3 vs QwQ-32B
- Grok 3 vs Llama 3.3 Nemotron 49b Super v1
- Grok 3 vs Step 1o Turbo 202506
- Grok 3 vs GPT-6 Astra
- Grok 3 vs Claude Fable 5.1
- Grok 3 vs Gemini 3.8 Flash
- Grok 3 vs Kimi K3
- Grok 3 vs Qwen3.8 Max
- Grok 3 vs GLM-5.3
- Grok 3 vs Muse Spark 1.3
- Grok 3 vs DeepSeek V4 Pro
Other xAI models
- Grok 4.656.9
- Grok 4.555.0
- Grok 4.753.1
- Grok 4.20 (Non-Reasoning)48.6
- Grok 448.1
- Grok 4.20 Multi-Agent46.2
- Grok 4.343.8
- Grok 4.141.5
Frequently asked questions
How good is Grok 3?
Grok 3 by xAI ranks 157th of 354 ranked models on the Noometry Index as of October 2026, with a score of 39.9. Its strongest category is instruction following, where it ranks 73rd.
Is Grok 3 open source?
No. Grok 3 is proprietary and available only through xAI's API and partner platforms.
How fast is Grok 3?
Grok 3 generated about 42 output tokens per second in the Kagi LLM Benchmark's timed runs. Speed varies by provider, load and reasoning effort.
What are Grok 3's strengths and weaknesses?
Relative to other ranked models, Grok 3 places best in instruction following, knowledge, multilingual and lowest in reasoning, long context, agentic & tool use.
What is Grok 3 best at?
Its best category is instruction following, where it ranks 73rd on Noometry.