Model comparison

Grok 4.6 vs Inkling

Grok 4.6 is the stronger model overall, scoring 56.9 to 44.1 on the Noometry Index.

Last verified . 38 shared benchmarks.

Grok 4.6 xAI

56.9

Rank #21 Confirmed

Inkling Thinking Machines Lab

44.1

Rank #80 Confirmed

Summary

  • They share 38 benchmarks with published results for both. Grok 4.6 scores higher in 7 categories and Inkling in 2 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in math, where Grok 4.6 leads 67.0 to 31.3.
  • The biggest single-benchmark swing is ProofBench: 51% for Grok 4.6 and 0% for Inkling.
  • Inkling is cheaper at $1.87 / $4.68 per million input/output tokens, against $2 / $6 for Grok 4.6.
  • Grok 4.6 accepts more context: 500K tokens versus 66K.
  • Inkling has downloadable open weights; the other is API-only.

Side by side

Grok 4.6 and Inkling specifications
Grok 4.6Inkling
ProviderxAIThinking Machines Lab
Noometry Index56.944.1
Released2026-08-122026-07-15
WeightsProprietaryOpen
Context window500K66K
Max output500K66K
Input $ / M tokens$2$1.87
Output $ / M tokens$6$4.68
Results tracked4941

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Grok 4.6 leads

Grok 4.6: 58.5 (#16), Inkling: 34.5 (#234)

Coding benchmarks
BenchmarkGrok 4.6Inkling
FrontierCode48%14%
LMArena WebDev16171413
FrontierSWE25.3%4.1%
SciCode56.5%47%
WeirdML67.3%32.3%
LMArena Coding14651464
ALE-Bench1,508946
DeepSWE67.5%—
CursorBench41.4%—

Agentic & Tool Use Grok 4.6 leads

Grok 4.6: 39.4 (#27), Inkling: 29.6 (#85)

Agentic & Tool Use benchmarks
BenchmarkGrok 4.6Inkling
APEX-Agents65.3%33.8%
τ²-bench Banking—25%
GDP.pdf17.2%—
Vending-Bench 29,047—

Reasoning Grok 4.6 leads

Grok 4.6: 61.4 (#20), Inkling: 40.4 (#56)

Reasoning benchmarks
BenchmarkGrok 4.6Inkling
ARC-AGI-267.1%36.5%
SimpleBench75.9%50%
ARC-AGI-187.5%79.5%
CritPt19.7%5.4%
Chess Puzzles40%21%
LMArena Hard Prompts14471451
DTBench97.3%87.5%
LMCA48.5%37.6%
Epoch Capabilities Index156.44148.54
NYT Connections (extended)80%—
EBR-Bench30.5%—
Mystery Game Puzzles34%—

Math Grok 4.6 leads

Grok 4.6: 67.0 (#24), Inkling: 31.3 (#225)

Math benchmarks
BenchmarkGrok 4.6Inkling
FrontierMath (Tiers 1-3)66%33.3%
FrontierMath Tier 431.7%4.9%
OTIS Mock AIME 2024-202599.2%88.9%
ProofBench51%0%
LMArena Math14231479

Knowledge Grok 4.6 leads

Grok 4.6: 63.3 (#20), Inkling: 55.1 (#49)

Knowledge benchmarks
BenchmarkGrok 4.6Inkling
GPQA Diamond94%88.3%
SimpleQA Verified49.3%40.3%
LMArena Expert14671465

Multimodal Not comparable

Grok 4.6: 43.6 (#23), Inkling: —

Multimodal benchmarks
BenchmarkGrok 4.6Inkling
LMArena Vision1263—
Blueprint-Bench 233.2%—
Furniture Assembly40%—
LMArena Document1452—

Multilingual Inkling leads

Grok 4.6: 53.0 (#74), Inkling: 54.0 (#52)

Multilingual benchmarks
BenchmarkGrok 4.6Inkling
LMArena Non-English14201434
LMArena Chinese14801490
LMArena French14611458
LMArena German14311446
LMArena Japanese13761429
LMArena Korean13971404
LMArena Russian14221429
LMArena Spanish14041448

Instruction Following Too close to call

Grok 4.6: 75.4 (#63), Inkling: 75.1 (#71)

Instruction Following benchmarks
BenchmarkGrok 4.6Inkling
LMArena Instruction Following14311426

Long Context Too close to call

Grok 4.6: 44.5 (#66), Inkling: 43.8 (#86)

Long Context benchmarks
BenchmarkGrok 4.6Inkling
LMArena Longer Query14541434

Writing & Preference Inkling leads

Grok 4.6: 62.3 (#80), Inkling: 65.2 (#51)

Writing & Preference benchmarks
BenchmarkGrok 4.6Inkling
LMArena Text14281441
LMArena Creative Writing14281387
LMArena Multi-Turn14251436
EQ-Bench Creative Writing—1611
EQ-Bench 4—1226

Frequently asked questions

Is Grok 4.6 better than Inkling?

Grok 4.6 is the stronger model overall, scoring 56.9 to 44.1 on the Noometry Index.

Which is cheaper, Grok 4.6 or Inkling?

Inkling is cheaper. It lists at $1.87 per million input tokens and $4.68 per million output tokens; Grok 4.6 lists at $2 and $6.

Is Grok 4.6 or Inkling better for coding?

Grok 4.6 scores higher on coding benchmarks: 58.5 versus 34.5 in the Noometry coding category.

Which has the bigger context window?

Grok 4.6 does, with 500K tokens against 66K.

How many benchmarks do Grok 4.6 and Inkling share?

38 benchmarks have published results for both models. Grok 4.6 has 49 scored results on Noometry and Inkling has 41.

Related comparisons

Go deeper