Model comparison

Grok 4.20 Multi-Agent vs Inkling-Small

Grok 4.20 Multi-Agent and Inkling-Small score almost the same on the Noometry Index (46.2 vs 46.5), so choose on price, context window or the category you care about most.

Last verified . 18 shared benchmarks.

Grok 4.20 Multi-Agent xAI

46.2

Rank #65 Confirmed

Inkling-Small Thinking Machines Lab

46.5

Rank #63 Confirmed

Summary

  • They share 18 benchmarks with published results for both. Grok 4.20 Multi-Agent scores higher in 6 categories and Inkling-Small in 3 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Inkling-Small leads 48.2 to 40.4.
  • Inkling-Small is cheaper at $0.45 / $1.20 per million input/output tokens, against $1.25 / $2.50 for Grok 4.20 Multi-Agent.
  • Grok 4.20 Multi-Agent accepts more context: 1M tokens versus 524K.
  • Inkling-Small has downloadable open weights; the other is API-only.

Side by side

Grok 4.20 Multi-Agent and Inkling-Small specifications
Grok 4.20 Multi-AgentInkling-Small
ProviderxAIThinking Machines Lab
Noometry Index46.246.5
Released2026-03-092026-07-15
WeightsProprietaryOpen
Context window1M524K
Max output30K1.05M
Input $ / M tokens$1.25$0.45
Output $ / M tokens$2.50$1.20
Results tracked2033

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

Grok 4.20 Multi-Agent: 43.0 (#92), Inkling-Small: 43.6 (#85)

Coding benchmarks
BenchmarkGrok 4.20 Multi-AgentInkling-Small
LMArena Coding14571451
LMArena WebDev—1409
SciCode—48.7%

Agentic & Tool Use Not comparable

Grok 4.20 Multi-Agent: —, Inkling-Small: —

Agentic & Tool Use benchmarks
BenchmarkGrok 4.20 Multi-AgentInkling-Small
LMArena Search1204—

Reasoning Grok 4.20 Multi-Agent leads

Grok 4.20 Multi-Agent: 43.9 (#48), Inkling-Small: 38.6 (#63)

Reasoning benchmarks
BenchmarkGrok 4.20 Multi-AgentInkling-Small
LMArena Hard Prompts14481423
ARC-AGI-2—40.1%
NYT Connections (extended)89.6%—
ARC-AGI-1—84%
CritPt—8.3%
Chess Puzzles—18%
Mystery Game Puzzles—6%
Epoch Capabilities Index—150.15

Math Inkling-Small leads

Grok 4.20 Multi-Agent: 39.4 (#104), Inkling-Small: 45.1 (#77)

Math benchmarks
BenchmarkGrok 4.20 Multi-AgentInkling-Small
LMArena Math14421459
FrontierMath (Tiers 1-3)—46.3%
FrontierMath Tier 4—17.1%
OTIS Mock AIME 2024-2025—90%
ProofBench—6%

Knowledge Inkling-Small leads

Grok 4.20 Multi-Agent: 40.4 (#119), Inkling-Small: 48.2 (#77)

Knowledge benchmarks
BenchmarkGrok 4.20 Multi-AgentInkling-Small
LMArena Expert14451442
GPQA Diamond—88.5%
SimpleQA Verified—19.1%

Multimodal Grok 4.20 Multi-Agent leads

Grok 4.20 Multi-Agent: 40.5 (#48), Inkling-Small: 39.1 (#62)

Multimodal benchmarks
BenchmarkGrok 4.20 Multi-AgentInkling-Small
LMArena Vision12591235

Multilingual Grok 4.20 Multi-Agent leads

Grok 4.20 Multi-Agent: 54.4 (#43), Inkling-Small: 51.7 (#104)

Multilingual benchmarks
BenchmarkGrok 4.20 Multi-AgentInkling-Small
LMArena Non-English14401402
LMArena Chinese14751465
LMArena French14661436
LMArena German14561405
LMArena Japanese14051405
LMArena Korean14161363
LMArena Russian14571391
LMArena Spanish14471428

Instruction Following Grok 4.20 Multi-Agent leads

Grok 4.20 Multi-Agent: 74.8 (#84), Inkling-Small: 73.8 (#114)

Instruction Following benchmarks
BenchmarkGrok 4.20 Multi-AgentInkling-Small
LMArena Instruction Following14201399

Long Context Too close to call

Grok 4.20 Multi-Agent: 43.7 (#88), Inkling-Small: 42.7 (#118)

Long Context benchmarks
BenchmarkGrok 4.20 Multi-AgentInkling-Small
LMArena Longer Query14311401

Writing & Preference Grok 4.20 Multi-Agent leads

Grok 4.20 Multi-Agent: 64.0 (#59), Inkling-Small: 59.6 (#107)

Writing & Preference benchmarks
BenchmarkGrok 4.20 Multi-AgentInkling-Small
LMArena Text14501414
LMArena Creative Writing14361331
LMArena Multi-Turn14521418
EQ-Bench Creative Writing—1491

Frequently asked questions

Is Grok 4.20 Multi-Agent better than Inkling-Small?

Grok 4.20 Multi-Agent and Inkling-Small score almost the same on the Noometry Index (46.2 vs 46.5), so choose on price, context window or the category you care about most.

Which is cheaper, Grok 4.20 Multi-Agent or Inkling-Small?

Inkling-Small is cheaper. It lists at $0.45 per million input tokens and $1.20 per million output tokens; Grok 4.20 Multi-Agent lists at $1.25 and $2.50.

Is Grok 4.20 Multi-Agent or Inkling-Small better for coding?

They score almost the same on coding (43.0 vs 43.6); test both on your own repository before choosing.

Which has the bigger context window?

Grok 4.20 Multi-Agent does, with 1M tokens against 524K.

How many benchmarks do Grok 4.20 Multi-Agent and Inkling-Small share?

18 benchmarks have published results for both models. Grok 4.20 Multi-Agent has 20 scored results on Noometry and Inkling-Small has 33.

Related comparisons

Go deeper