xAI, proprietary

Grok 4.20 (Non-Reasoning)

Grok 4.20 (Non-Reasoning) by xAI ranks 54th of 354 ranked models on the Noometry Index as of October 2026, with a score of 48.6. Its strongest category is reasoning, where it ranks 32nd. API pricing starts at $1.25 per million input tokens and $2.50 per million output tokens, with a 1M-token context window.

Last verified

Specifications

Noometry rank
#54 of 354
Index score
48.6
Evidence
Confirmed 46 results
Provider
xAI
Released
February 17, 2026
Weights
Proprietary
Reasoning
No
Context window
1M
Max output
30K
Input price
$1.25 / M
Output price
$2.50 / M
Blended price
$1.56 / M
Output speed
61 tokens/s Kagi
Value
#132 of 219
Knowledge cutoff
September 2025
Input
text, image, pdf

Category scores

Each category score combines every public result we have in that category.

Grok 4.20 (Non-Reasoning) category scores
  1. Coding 42.1
  2. Agentic & Tool Use 34.4
  3. Reasoning 52.3
  4. Math 48.2
  5. Knowledge 52.8
  6. Multimodal 33.3
  7. Multilingual 54.5
  8. Instruction Following 74.8
  9. Long Context 45.5
  10. Writing & Preference 65.7
Grok 4.20 (Non-Reasoning) category ranks
CategoryScoreRankResults
Coding42.1#1123
Agentic & Tool Use34.4#462
Reasoning52.3#329
Math48.2#655
Knowledge52.8#603
Multimodal33.3#982
Multilingual54.5#401
Instruction Following74.8#831
Long Context45.5#343
Writing & Preference65.7#444

Strengths and weaknesses

Categories where Grok 4.20 (Non-Reasoning) places highest and lowest among the models ranked in each, with its score against that category's median.

Strongest categories

Grok 4.20 (Non-Reasoning): strongest categories
CategoryScorevs medianRank
Reasoning52.3+28.7#32 of 350, top 10%
Long Context45.5+4.5#34 of 296, top 12%
Multilingual54.5+7.1#40 of 297, top 14%

Weakest categories

Grok 4.20 (Non-Reasoning): weakest categories
CategoryScorevs medianRank
Multimodal33.3−5.3#98 of 128, top 77%
Coding42.1+3.3#112 of 340, top 33%
Agentic & Tool Use34.4+4.1#46 of 154, top 30%

Closest competitors

The models ranked just above and below Grok 4.20 (Non-Reasoning). When scores are this close, price and speed are often the better way to choose.

Models ranked closest to Grok 4.20 (Non-Reasoning)
ModelRankScoreBlended $/MSpeed
Claude Sonnet 4.6#5050.3$6—Compare
Muse Spark 1.1#5149.9$2—Compare
Claude Haiku 5.5#5249.5$0.20—Compare
GPT-5.1#5349.0$3.44—Compare
MiMo-V2.6-Flash#5548.5$0.18—Compare
Grok 4#5648.1—1Compare
Kimi K2.5#5748.1$0.9066Compare
Step 5 Preview#5847.9$1.43—Compare

Sponsored placements are available on pages like this one. Advertise on Noometry

Benchmark results

Every published result we track, with its source. Bold rows are the ones used for ranking; where several exist we prefer independent runs over self-reported numbers.

Coding

Grok 4.20 (Non-Reasoning) Coding benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena WebDev1375#80 of 113, top 71%LMArena2026-10-08
WeirdML52.3%#46 of 119, top 39%Epoch AI
LMArena Coding1459#73 of 294, top 25%LMArena2026-10-08
LMArena Coding1447LMArena2026-10-08
ALE-Bench1,150#37 of 105, top 36%Epoch AI

Agentic & Tool Use

Grok 4.20 (Non-Reasoning) Agentic & Tool Use benchmark results
BenchmarkScorePositionSettingSourceDate
Terminal-Bench57.3%#14 of 41, top 35%Epoch AI
τ²-bench Banking18%#21 of 26, top 81%highτ²-bench2026-05-05
LMArena Search1189#18 of 32, top 57%LMArena2026-08-24
Vending-Bench 24,663#32 of 60, top 54%Epoch AI

Reasoning

Grok 4.20 (Non-Reasoning) Reasoning benchmark results
BenchmarkScorePositionSettingSourceDate
ARC-AGI-265.1%#24 of 83, top 29%Epoch AI
Kagi LLM Benchmark44.8%Kagi LLM Benchmark
Kagi LLM Benchmark75%#13 of 99, top 14%Kagi LLM Benchmark
NYT Connections (extended)83.7%Lech Mazur benchmarks
NYT Connections (extended)7.6%Lech Mazur benchmarks
NYT Connections (extended)85.4%#27 of 91, top 30%reasoningLech Mazur benchmarks
ARC-AGI-189.5%#27 of 83, top 33%Epoch AI
Chess Puzzles24%#46 of 129, top 36%Epoch AI2026-07-13
Thematic Generalization63.8%#10 of 23, top 44%reasoningLech Mazur benchmarks
LMArena Hard Prompts1451#59 of 297, top 20%LMArena2026-10-08
LMArena Hard Prompts1441LMArena2026-10-08
DTBench90.1%#39 of 151, top 26%Epoch AI
LMCA38.7%#49 of 125, top 40%Epoch AI
Epoch Capabilities Index151.98#45 of 213, top 22%Epoch AI2026-02-17
ForecastBench60.7Epoch AI
ForecastBench61.4#11 of 72, top 16%Epoch AI

Math

Grok 4.20 (Non-Reasoning) Math benchmark results
BenchmarkScorePositionSettingSourceDate
FrontierMath (Tiers 1-3)44.9%#51 of 81, top 63%Epoch AI2026-07-13
FrontierMath Tier 417.1%#46 of 63, top 74%Epoch AI2026-07-13
OTIS Mock AIME 2024-202592.2%#45 of 173, top 27%Epoch AI2026-07-13
ProofBench14%#59 of 77, top 77%Epoch AI
LMArena Math1433LMArena2026-10-08
LMArena Math1455#57 of 285, top 20%LMArena2026-10-08

Knowledge

Grok 4.20 (Non-Reasoning) Knowledge benchmark results
BenchmarkScorePositionSettingSourceDate
GPQA Diamond89.3%#44 of 186, top 24%Epoch AI2026-07-13
SimpleQA Verified30.2%#58 of 77, top 76%Epoch AI2026-08-27
LMArena Expert1428LMArena2026-10-08
LMArena Expert1439#87 of 273, top 32%LMArena2026-10-08

Multimodal

Grok 4.20 (Non-Reasoning) Multimodal benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Vision1263#45 of 122, top 37%LMArena2026-10-09
Blueprint-Bench 20%#30 of 31, top 97%Epoch AI
LMArena Document1416#33 of 38, top 87%LMArena2026-09-13

Multilingual

Grok 4.20 (Non-Reasoning) Multilingual benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Non-English1441#40 of 297, top 14%LMArena2026-10-08
LMArena Non-English1433LMArena2026-10-08
LMArena Chinese1481#63 of 285, top 23%LMArena2026-10-08
LMArena Chinese1476LMArena2026-10-08
LMArena French1476#30 of 223, top 14%LMArena2026-10-08
LMArena French1455LMArena2026-10-08
LMArena German1440LMArena2026-10-08
LMArena German1465#32 of 231, top 14%LMArena2026-10-08
LMArena Japanese1415LMArena2026-10-08
LMArena Japanese1449#26 of 211, top 13%LMArena2026-10-08
LMArena Korean1417#37 of 213, top 18%LMArena2026-10-08
LMArena Korean1416LMArena2026-10-08
LMArena Russian1458#35 of 283, top 13%LMArena2026-10-08
LMArena Russian1450LMArena2026-10-08
LMArena Spanish1438LMArena2026-10-08
LMArena Spanish1443#62 of 226, top 28%LMArena2026-10-08

Instruction Following

Grok 4.20 (Non-Reasoning) Instruction Following benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Instruction Following1420#73 of 298, top 25%LMArena2026-10-08
LMArena Instruction Following1415LMArena2026-10-08

Long Context

Grok 4.20 (Non-Reasoning) Long Context benchmark results
BenchmarkScorePositionSettingSourceDate
CL-bench22.2%#3 of 19, top 16%Epoch AI
CL-bench Life11.9%#9 of 13, top 70%Epoch AI
LMArena Longer Query1427LMArena2026-10-08
LMArena Longer Query1437#70 of 291, top 25%LMArena2026-10-08

Writing & Preference

Grok 4.20 (Non-Reasoning) Writing & Preference benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Text1451#43 of 297, top 15%LMArena2026-10-08
LMArena Text1445LMArena2026-10-08
LMArena Creative Writing1438#41 of 295, top 14%LMArena2026-10-08
LMArena Creative Writing1432LMArena2026-10-08
EQ-Bench Creative Writing1574#47 of 115, top 41%EQ-Bench
LMArena Multi-Turn1456#45 of 295, top 16%LMArena2026-10-08
LMArena Multi-Turn1451LMArena2026-10-08

API pricing by provider

Grok 4.20 (Non-Reasoning) API prices
RouteInput $/MOutput $/MCached input $/MChecked
azure$2$6—2026-10-10
openrouter$1.25$2.50$0.202026-10-10
vertex$1.25$2.50$0.202026-10-10
xai$1.25$2.50$0.202026-10-10

Compare Grok 4.20 (Non-Reasoning)

Other xAI models

Frequently asked questions

How good is Grok 4.20 (Non-Reasoning)?

Grok 4.20 (Non-Reasoning) by xAI ranks 54th of 354 ranked models on the Noometry Index as of October 2026, with a score of 48.6. Its strongest category is reasoning, where it ranks 32nd. API pricing starts at $1.25 per million input tokens and $2.50 per million output tokens, with a 1M-token context window.

How much does Grok 4.20 (Non-Reasoning) cost?

Grok 4.20 (Non-Reasoning) costs $1.25 per million input tokens and $2.50 per million output tokens on xAI's own API, with cached input at $0.20.

What is Grok 4.20 (Non-Reasoning)'s context window?

Grok 4.20 (Non-Reasoning) accepts up to 1M tokens of input and can write up to 30K tokens in one response.

Is Grok 4.20 (Non-Reasoning) open source?

No. Grok 4.20 (Non-Reasoning) is proprietary and available only through xAI's API and partner platforms.

How fast is Grok 4.20 (Non-Reasoning)?

Grok 4.20 (Non-Reasoning) generated about 61 output tokens per second in the Kagi LLM Benchmark's timed runs. Speed varies by provider, load and reasoning effort.

What are Grok 4.20 (Non-Reasoning)'s strengths and weaknesses?

Relative to other ranked models, Grok 4.20 (Non-Reasoning) places best in reasoning, long context, multilingual and lowest in multimodal, coding, agentic & tool use.

What is Grok 4.20 (Non-Reasoning) best at?

Its best category is reasoning, where it ranks 32nd on Noometry.