xAI, proprietary

Grok 4.20 Multi-Agent

Grok 4.20 Multi-Agent by xAI ranks 65th of 354 ranked models on the Noometry Index as of October 2026, with a score of 46.2. Its strongest category is multilingual, where it ranks 43rd. API pricing starts at $1.25 per million input tokens and $2.50 per million output tokens, with a 1M-token context window.

Last verified

Specifications

Noometry rank
#65 of 354
Index score
46.2
Evidence
Confirmed 20 results
Provider
xAI
Released
March 9, 2026
Weights
Proprietary
Reasoning
Yes
Context window
1M
Max output
30K
Input price
$1.25 / M
Output price
$2.50 / M
Blended price
$1.56 / M
Output speed
Not measured
Value
#135 of 219
Knowledge cutoff
Unknown
Input
text, image, pdf

Category scores

Each category score combines every public result we have in that category.

Grok 4.20 Multi-Agent category scores
  1. Coding 43.0
  2. Reasoning 43.9
  3. Math 39.4
  4. Knowledge 40.4
  5. Multimodal 40.5
  6. Multilingual 54.4
  7. Instruction Following 74.8
  8. Long Context 43.7
  9. Writing & Preference 64.0
Grok 4.20 Multi-Agent category ranks
CategoryScoreRankResults
Coding43.0#921
Reasoning43.9#482
Math39.4#1041
Knowledge40.4#1191
Multimodal40.5#481
Multilingual54.4#431
Instruction Following74.8#841
Long Context43.7#881
Writing & Preference64.0#593

Strengths and weaknesses

Categories where Grok 4.20 Multi-Agent places highest and lowest among the models ranked in each, with its score against that category's median.

Strongest categories

Grok 4.20 Multi-Agent: strongest categories
CategoryScorevs medianRank
Reasoning43.9+20.3#48 of 350, top 14%
Multilingual54.4+7.0#43 of 297, top 15%
Writing & Preference64.0+10.2#59 of 312, top 19%

Weakest categories

Grok 4.20 Multi-Agent: weakest categories
CategoryScorevs medianRank
Knowledge40.4+3.0#119 of 314, top 38%
Multimodal40.5+2.0#48 of 128, top 38%
Math39.4+2.8#104 of 327, top 32%

Closest competitors

The models ranked just above and below Grok 4.20 Multi-Agent. When scores are this close, price and speed are often the better way to choose.

Models ranked closest to Grok 4.20 Multi-Agent
ModelRankScoreBlended $/MSpeed
o3#6147.5$3.503Compare
Qwen3.6 Plus#6247.5$1.13—Compare
Inkling-Small#6346.5$0.64—Compare
GPT-5 Pro#6446.4$41.255Compare
GLM-5#6646.1$1.5523Compare
Qwen3.5 397B-A17B#6746.0$1.359Compare
Qwen3.8 27B#6846.0$1.11—Compare
GPT-5.3 Codex#6945.8$4.81—Compare

Sponsored placements are available on pages like this one. Advertise on Noometry

Benchmark results

Every published result we track, with its source. Bold rows are the ones used for ranking; where several exist we prefer independent runs over self-reported numbers.

Coding

Grok 4.20 Multi-Agent Coding benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Coding1457#74 of 294, top 26%LMArena2026-10-08

Agentic & Tool Use

Grok 4.20 Multi-Agent Agentic & Tool Use benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Search1204#13 of 32, top 41%LMArena2026-08-24

Reasoning

Grok 4.20 Multi-Agent Reasoning benchmark results
BenchmarkScorePositionSettingSourceDate
NYT Connections (extended)89.6%#21 of 91, top 24%Lech Mazur benchmarks
LMArena Hard Prompts1448#64 of 297, top 22%LMArena2026-10-08

Math

Grok 4.20 Multi-Agent Math benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Math1442#67 of 285, top 24%LMArena2026-10-08

Knowledge

Grok 4.20 Multi-Agent Knowledge benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Expert1445#75 of 273, top 28%LMArena2026-10-08

Multimodal

Grok 4.20 Multi-Agent Multimodal benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Vision1259#49 of 122, top 41%LMArena2026-10-09

Multilingual

Grok 4.20 Multi-Agent Multilingual benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Non-English1440#43 of 297, top 15%LMArena2026-10-08
LMArena Chinese1475#71 of 285, top 25%LMArena2026-10-08
LMArena French1466#43 of 223, top 20%LMArena2026-10-08
LMArena German1456#39 of 231, top 17%LMArena2026-10-08
LMArena Japanese1405#55 of 211, top 27%LMArena2026-10-08
LMArena Korean1416#38 of 213, top 18%LMArena2026-10-08
LMArena Russian1457#36 of 283, top 13%LMArena2026-10-08
LMArena Spanish1447#57 of 226, top 26%LMArena2026-10-08

Instruction Following

Grok 4.20 Multi-Agent Instruction Following benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Instruction Following1420#74 of 298, top 25%LMArena2026-10-08

Long Context

Grok 4.20 Multi-Agent Long Context benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Longer Query1431#77 of 291, top 27%LMArena2026-10-08

Writing & Preference

Grok 4.20 Multi-Agent Writing & Preference benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Text1450#44 of 297, top 15%LMArena2026-10-08
LMArena Creative Writing1436#43 of 295, top 15%LMArena2026-10-08
LMArena Multi-Turn1452#52 of 295, top 18%LMArena2026-10-08

API pricing by provider

Grok 4.20 Multi-Agent API prices
RouteInput $/MOutput $/MCached input $/MChecked
openrouter$1.25$2.50$0.202026-10-10
xai$1.25$2.50$0.202026-10-10

Compare Grok 4.20 Multi-Agent

Other xAI models

Frequently asked questions

How good is Grok 4.20 Multi-Agent?

Grok 4.20 Multi-Agent by xAI ranks 65th of 354 ranked models on the Noometry Index as of October 2026, with a score of 46.2. Its strongest category is multilingual, where it ranks 43rd. API pricing starts at $1.25 per million input tokens and $2.50 per million output tokens, with a 1M-token context window.

How much does Grok 4.20 Multi-Agent cost?

Grok 4.20 Multi-Agent costs $1.25 per million input tokens and $2.50 per million output tokens on xAI's own API, with cached input at $0.20.

What is Grok 4.20 Multi-Agent's context window?

Grok 4.20 Multi-Agent accepts up to 1M tokens of input and can write up to 30K tokens in one response.

Is Grok 4.20 Multi-Agent open source?

No. Grok 4.20 Multi-Agent is proprietary and available only through xAI's API and partner platforms.

What are Grok 4.20 Multi-Agent's strengths and weaknesses?

Relative to other ranked models, Grok 4.20 Multi-Agent places best in reasoning, multilingual, writing & preference and lowest in knowledge, multimodal, math.

What is Grok 4.20 Multi-Agent best at?

Its best category is multilingual, where it ranks 43rd on Noometry.