Model comparison

Grok 4.1 vs Inkling

Inkling is the stronger model overall, scoring 44.1 to 41.5 on the Noometry Index.

Last verified . 18 shared benchmarks.

Grok 4.1 xAI

41.5

Rank #134 Confirmed

Inkling Thinking Machines Lab

44.1

Rank #80 Confirmed

Summary

  • They share 18 benchmarks with published results for both. Grok 4.1 scores higher in 2 categories and Inkling in 7 categories; 6 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Inkling leads 55.1 to 39.5.
  • Inkling has downloadable open weights; the other is API-only.

Side by side

Grok 4.1 and Inkling specifications
Grok 4.1Inkling
ProviderxAIThinking Machines Lab
Noometry Index41.544.1
Released2025-11-172026-07-15
WeightsProprietaryOpen
Context window—66K
Max output—66K
Input $ / M tokens—$1.87
Output $ / M tokens—$4.68
Results tracked1941

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

Grok 4.1: 33.7 (#253), Inkling: 34.5 (#234)

Coding benchmarks
BenchmarkGrok 4.1Inkling
LMArena WebDev12141413
LMArena Coding14451464
FrontierCode—14%
FrontierSWE—4.1%
SciCode—47%
WeirdML—32.3%
ALE-Bench—946

Agentic & Tool Use Grok 4.1 leads

Grok 4.1: 34.1 (#49), Inkling: 29.6 (#85)

Agentic & Tool Use benchmarks
BenchmarkGrok 4.1Inkling
APEX-Agents—33.8%
τ²-bench Banking—25%
Cybench39%—

Reasoning Inkling leads

Grok 4.1: 29.5 (#91), Inkling: 40.4 (#56)

Reasoning benchmarks
BenchmarkGrok 4.1Inkling
LMArena Hard Prompts14351451
ARC-AGI-2—36.5%
SimpleBench—50%
ARC-AGI-1—79.5%
CritPt—5.4%
Chess Puzzles—21%
DTBench—87.5%
LMCA—37.6%
Epoch Capabilities Index—148.54

Math Grok 4.1 leads

Grok 4.1: 38.9 (#120), Inkling: 31.3 (#225)

Math benchmarks
BenchmarkGrok 4.1Inkling
LMArena Math14221479
FrontierMath (Tiers 1-3)—33.3%
FrontierMath Tier 4—4.9%
OTIS Mock AIME 2024-2025—88.9%
ProofBench—0%

Knowledge Inkling leads

Grok 4.1: 39.5 (#133), Inkling: 55.1 (#49)

Knowledge benchmarks
BenchmarkGrok 4.1Inkling
LMArena Expert14171465
GPQA Diamond—88.3%
SimpleQA Verified—40.3%

Multilingual Too close to call

Grok 4.1: 53.4 (#68), Inkling: 54.0 (#52)

Multilingual benchmarks
BenchmarkGrok 4.1Inkling
LMArena Non-English14251434
LMArena Chinese14651490
LMArena French14481458
LMArena German14461446
LMArena Japanese13971429
LMArena Korean14071404
LMArena Russian14341429
LMArena Spanish14381448

Instruction Following Inkling leads

Grok 4.1: 73.8 (#111), Inkling: 75.1 (#71)

Instruction Following benchmarks
BenchmarkGrok 4.1Inkling
LMArena Instruction Following14001426

Long Context Too close to call

Grok 4.1: 43.2 (#100), Inkling: 43.8 (#86)

Long Context benchmarks
BenchmarkGrok 4.1Inkling
LMArena Longer Query14161434

Writing & Preference Inkling leads

Grok 4.1: 62.4 (#75), Inkling: 65.2 (#51)

Writing & Preference benchmarks
BenchmarkGrok 4.1Inkling
LMArena Text14371441
LMArena Creative Writing14111387
LMArena Multi-Turn14371436
EQ-Bench Creative Writing—1611
EQ-Bench 4—1226

Frequently asked questions

Is Grok 4.1 better than Inkling?

Inkling is the stronger model overall, scoring 44.1 to 41.5 on the Noometry Index.

Is Grok 4.1 or Inkling better for coding?

They score almost the same on coding (33.7 vs 34.5); test both on your own repository before choosing.

How many benchmarks do Grok 4.1 and Inkling share?

18 benchmarks have published results for both models. Grok 4.1 has 19 scored results on Noometry and Inkling has 41.

Related comparisons

Go deeper