xAI, proprietary
Grok 4.20 (Non-Reasoning)
Grok 4.20 (Non-Reasoning) by xAI ranks 54th of 354 ranked models on the Noometry Index as of October 2026, with a score of 48.6. Its strongest category is reasoning, where it ranks 32nd. API pricing starts at $1.25 per million input tokens and $2.50 per million output tokens, with a 1M-token context window.
Last verified
Specifications
- Noometry rank
- #54 of 354
- Index score
- 48.6
- Evidence
- Confirmed 46 results
- Provider
- xAI
- Released
- February 17, 2026
- Weights
- Proprietary
- Reasoning
- No
- Context window
- 1M
- Max output
- 30K
- Input price
- $1.25 / M
- Output price
- $2.50 / M
- Blended price
- $1.56 / M
- Output speed
- 61 tokens/s Kagi
- Value
- #132 of 219
- Knowledge cutoff
- September 2025
- Input
- text, image, pdf
Category scores
Each category score combines every public result we have in that category.
- Coding 42.1
- Agentic & Tool Use 34.4
- Reasoning 52.3
- Math 48.2
- Knowledge 52.8
- Multimodal 33.3
- Multilingual 54.5
- Instruction Following 74.8
- Long Context 45.5
- Writing & Preference 65.7
| Category | Score | Rank | Results |
|---|---|---|---|
| Coding | 42.1 | #112 | 3 |
| Agentic & Tool Use | 34.4 | #46 | 2 |
| Reasoning | 52.3 | #32 | 9 |
| Math | 48.2 | #65 | 5 |
| Knowledge | 52.8 | #60 | 3 |
| Multimodal | 33.3 | #98 | 2 |
| Multilingual | 54.5 | #40 | 1 |
| Instruction Following | 74.8 | #83 | 1 |
| Long Context | 45.5 | #34 | 3 |
| Writing & Preference | 65.7 | #44 | 4 |
Strengths and weaknesses
Categories where Grok 4.20 (Non-Reasoning) places highest and lowest among the models ranked in each, with its score against that category's median.
Strongest categories
| Category | Score | vs median | Rank |
|---|---|---|---|
| Reasoning | 52.3 | +28.7 | #32 of 350, top 10% |
| Long Context | 45.5 | +4.5 | #34 of 296, top 12% |
| Multilingual | 54.5 | +7.1 | #40 of 297, top 14% |
Weakest categories
| Category | Score | vs median | Rank |
|---|---|---|---|
| Multimodal | 33.3 | −5.3 | #98 of 128, top 77% |
| Coding | 42.1 | +3.3 | #112 of 340, top 33% |
| Agentic & Tool Use | 34.4 | +4.1 | #46 of 154, top 30% |
Closest competitors
The models ranked just above and below Grok 4.20 (Non-Reasoning). When scores are this close, price and speed are often the better way to choose.
| Model | Rank | Score | Blended $/M | Speed | |
|---|---|---|---|---|---|
| Claude Sonnet 4.6 | #50 | 50.3 | $6 | — | Compare |
| Muse Spark 1.1 | #51 | 49.9 | $2 | — | Compare |
| Claude Haiku 5.5 | #52 | 49.5 | $0.20 | — | Compare |
| GPT-5.1 | #53 | 49.0 | $3.44 | — | Compare |
| MiMo-V2.6-Flash | #55 | 48.5 | $0.18 | — | Compare |
| Grok 4 | #56 | 48.1 | — | 1 | Compare |
| Kimi K2.5 | #57 | 48.1 | $0.90 | 66 | Compare |
| Step 5 Preview | #58 | 47.9 | $1.43 | — | Compare |
Sponsored placements are available on pages like this one. Advertise on Noometry
Benchmark results
Every published result we track, with its source. Bold rows are the ones used for ranking; where several exist we prefer independent runs over self-reported numbers.
Coding
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena WebDev | 1375 | #80 of 113, top 71% | LMArena | 2026-10-08 | |
| WeirdML | 52.3% | #46 of 119, top 39% | Epoch AI | ||
| LMArena Coding | 1459 | #73 of 294, top 25% | LMArena | 2026-10-08 | |
| LMArena Coding | 1447 | LMArena | 2026-10-08 | ||
| ALE-Bench | 1,150 | #37 of 105, top 36% | Epoch AI |
Agentic & Tool Use
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| Terminal-Bench | 57.3% | #14 of 41, top 35% | Epoch AI | ||
| τ²-bench Banking | 18% | #21 of 26, top 81% | high | τ²-bench | 2026-05-05 |
| LMArena Search | 1189 | #18 of 32, top 57% | LMArena | 2026-08-24 | |
| Vending-Bench 2 | 4,663 | #32 of 60, top 54% | Epoch AI |
Reasoning
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| ARC-AGI-2 | 65.1% | #24 of 83, top 29% | Epoch AI | ||
| Kagi LLM Benchmark | 44.8% | Kagi LLM Benchmark | |||
| Kagi LLM Benchmark | 75% | #13 of 99, top 14% | Kagi LLM Benchmark | ||
| NYT Connections (extended) | 83.7% | Lech Mazur benchmarks | |||
| NYT Connections (extended) | 7.6% | Lech Mazur benchmarks | |||
| NYT Connections (extended) | 85.4% | #27 of 91, top 30% | reasoning | Lech Mazur benchmarks | |
| ARC-AGI-1 | 89.5% | #27 of 83, top 33% | Epoch AI | ||
| Chess Puzzles | 24% | #46 of 129, top 36% | Epoch AI | 2026-07-13 | |
| Thematic Generalization | 63.8% | #10 of 23, top 44% | reasoning | Lech Mazur benchmarks | |
| LMArena Hard Prompts | 1451 | #59 of 297, top 20% | LMArena | 2026-10-08 | |
| LMArena Hard Prompts | 1441 | LMArena | 2026-10-08 | ||
| DTBench | 90.1% | #39 of 151, top 26% | Epoch AI | ||
| LMCA | 38.7% | #49 of 125, top 40% | Epoch AI | ||
| Epoch Capabilities Index | 151.98 | #45 of 213, top 22% | Epoch AI | 2026-02-17 | |
| ForecastBench | 60.7 | Epoch AI | |||
| ForecastBench | 61.4 | #11 of 72, top 16% | Epoch AI |
Math
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| FrontierMath (Tiers 1-3) | 44.9% | #51 of 81, top 63% | Epoch AI | 2026-07-13 | |
| FrontierMath Tier 4 | 17.1% | #46 of 63, top 74% | Epoch AI | 2026-07-13 | |
| OTIS Mock AIME 2024-2025 | 92.2% | #45 of 173, top 27% | Epoch AI | 2026-07-13 | |
| ProofBench | 14% | #59 of 77, top 77% | Epoch AI | ||
| LMArena Math | 1433 | LMArena | 2026-10-08 | ||
| LMArena Math | 1455 | #57 of 285, top 20% | LMArena | 2026-10-08 |
Knowledge
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| GPQA Diamond | 89.3% | #44 of 186, top 24% | Epoch AI | 2026-07-13 | |
| SimpleQA Verified | 30.2% | #58 of 77, top 76% | Epoch AI | 2026-08-27 | |
| LMArena Expert | 1428 | LMArena | 2026-10-08 | ||
| LMArena Expert | 1439 | #87 of 273, top 32% | LMArena | 2026-10-08 |
Multimodal
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Vision | 1263 | #45 of 122, top 37% | LMArena | 2026-10-09 | |
| Blueprint-Bench 2 | 0% | #30 of 31, top 97% | Epoch AI | ||
| LMArena Document | 1416 | #33 of 38, top 87% | LMArena | 2026-09-13 |
Multilingual
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Non-English | 1441 | #40 of 297, top 14% | LMArena | 2026-10-08 | |
| LMArena Non-English | 1433 | LMArena | 2026-10-08 | ||
| LMArena Chinese | 1481 | #63 of 285, top 23% | LMArena | 2026-10-08 | |
| LMArena Chinese | 1476 | LMArena | 2026-10-08 | ||
| LMArena French | 1476 | #30 of 223, top 14% | LMArena | 2026-10-08 | |
| LMArena French | 1455 | LMArena | 2026-10-08 | ||
| LMArena German | 1440 | LMArena | 2026-10-08 | ||
| LMArena German | 1465 | #32 of 231, top 14% | LMArena | 2026-10-08 | |
| LMArena Japanese | 1415 | LMArena | 2026-10-08 | ||
| LMArena Japanese | 1449 | #26 of 211, top 13% | LMArena | 2026-10-08 | |
| LMArena Korean | 1417 | #37 of 213, top 18% | LMArena | 2026-10-08 | |
| LMArena Korean | 1416 | LMArena | 2026-10-08 | ||
| LMArena Russian | 1458 | #35 of 283, top 13% | LMArena | 2026-10-08 | |
| LMArena Russian | 1450 | LMArena | 2026-10-08 | ||
| LMArena Spanish | 1438 | LMArena | 2026-10-08 | ||
| LMArena Spanish | 1443 | #62 of 226, top 28% | LMArena | 2026-10-08 |
Instruction Following
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Instruction Following | 1420 | #73 of 298, top 25% | LMArena | 2026-10-08 | |
| LMArena Instruction Following | 1415 | LMArena | 2026-10-08 |
Long Context
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| CL-bench | 22.2% | #3 of 19, top 16% | Epoch AI | ||
| CL-bench Life | 11.9% | #9 of 13, top 70% | Epoch AI | ||
| LMArena Longer Query | 1427 | LMArena | 2026-10-08 | ||
| LMArena Longer Query | 1437 | #70 of 291, top 25% | LMArena | 2026-10-08 |
Writing & Preference
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Text | 1451 | #43 of 297, top 15% | LMArena | 2026-10-08 | |
| LMArena Text | 1445 | LMArena | 2026-10-08 | ||
| LMArena Creative Writing | 1438 | #41 of 295, top 14% | LMArena | 2026-10-08 | |
| LMArena Creative Writing | 1432 | LMArena | 2026-10-08 | ||
| EQ-Bench Creative Writing | 1574 | #47 of 115, top 41% | EQ-Bench | ||
| LMArena Multi-Turn | 1456 | #45 of 295, top 16% | LMArena | 2026-10-08 | |
| LMArena Multi-Turn | 1451 | LMArena | 2026-10-08 |
API pricing by provider
| Route | Input $/M | Output $/M | Cached input $/M | Checked |
|---|---|---|---|---|
| azure | $2 | $6 | — | 2026-10-10 |
| openrouter | $1.25 | $2.50 | $0.20 | 2026-10-10 |
| vertex | $1.25 | $2.50 | $0.20 | 2026-10-10 |
| xai | $1.25 | $2.50 | $0.20 | 2026-10-10 |
Compare Grok 4.20 (Non-Reasoning)
- Grok 4.20 (Non-Reasoning) vs GPT-5.1
- Grok 4.20 (Non-Reasoning) vs MiMo-V2.6-Flash
- Grok 4.20 (Non-Reasoning) vs Claude Haiku 5.5
- Grok 4.20 (Non-Reasoning) vs Grok 4
- Grok 4.20 (Non-Reasoning) vs Muse Spark 1.1
- Grok 4.20 (Non-Reasoning) vs Kimi K2.5
- Grok 4.20 (Non-Reasoning) vs GPT-6 Astra
- Grok 4.20 (Non-Reasoning) vs Claude Fable 5.1
- Grok 4.20 (Non-Reasoning) vs Gemini 3.8 Flash
- Grok 4.20 (Non-Reasoning) vs Kimi K3
- Grok 4.20 (Non-Reasoning) vs Qwen3.8 Max
- Grok 4.20 (Non-Reasoning) vs GLM-5.3
- Grok 4.20 (Non-Reasoning) vs Muse Spark 1.3
- Grok 4.20 (Non-Reasoning) vs DeepSeek V4 Pro
Other xAI models
- Grok 4.656.9
- Grok 4.555.0
- Grok 4.753.1
- Grok 448.1
- Grok 4.20 Multi-Agent46.2
- Grok 4.343.8
- Grok 4.141.5
- Grok 4.1 Fast41.4
Frequently asked questions
How good is Grok 4.20 (Non-Reasoning)?
Grok 4.20 (Non-Reasoning) by xAI ranks 54th of 354 ranked models on the Noometry Index as of October 2026, with a score of 48.6. Its strongest category is reasoning, where it ranks 32nd. API pricing starts at $1.25 per million input tokens and $2.50 per million output tokens, with a 1M-token context window.
How much does Grok 4.20 (Non-Reasoning) cost?
Grok 4.20 (Non-Reasoning) costs $1.25 per million input tokens and $2.50 per million output tokens on xAI's own API, with cached input at $0.20.
What is Grok 4.20 (Non-Reasoning)'s context window?
Grok 4.20 (Non-Reasoning) accepts up to 1M tokens of input and can write up to 30K tokens in one response.
Is Grok 4.20 (Non-Reasoning) open source?
No. Grok 4.20 (Non-Reasoning) is proprietary and available only through xAI's API and partner platforms.
How fast is Grok 4.20 (Non-Reasoning)?
Grok 4.20 (Non-Reasoning) generated about 61 output tokens per second in the Kagi LLM Benchmark's timed runs. Speed varies by provider, load and reasoning effort.
What are Grok 4.20 (Non-Reasoning)'s strengths and weaknesses?
Relative to other ranked models, Grok 4.20 (Non-Reasoning) places best in reasoning, long context, multilingual and lowest in multimodal, coding, agentic & tool use.
What is Grok 4.20 (Non-Reasoning) best at?
Its best category is reasoning, where it ranks 32nd on Noometry.