Model comparison

Grok 4.7 vs Inkling

Grok 4.7 is the stronger model overall, scoring 53.1 to 44.1 on the Noometry Index.

Last verified . 31 shared benchmarks.

Grok 4.7 xAI

53.1

Rank #37 Confirmed

Inkling Thinking Machines Lab

44.1

Rank #80 Confirmed

Summary

  • They share 31 benchmarks with published results for both. Grok 4.7 scores higher in 6 categories and Inkling in 3 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in math, where Grok 4.7 leads 57.8 to 31.3.
  • The biggest single-benchmark swing is ProofBench: 34% for Grok 4.7 and 0% for Inkling.
  • Inkling is cheaper at $1.87 / $4.68 per million input/output tokens, against $2 / $6 for Grok 4.7.
  • Grok 4.7 accepts more context: 500K tokens versus 66K.
  • Inkling has downloadable open weights; the other is API-only.

Side by side

Grok 4.7 and Inkling specifications
Grok 4.7Inkling
ProviderxAIThinking Machines Lab
Noometry Index53.144.1
Released2026-09-212026-07-15
WeightsProprietaryOpen
Context window500K66K
Max output500K66K
Input $ / M tokens$2$1.87
Output $ / M tokens$6$4.68
Results tracked3941

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Grok 4.7 leads

Grok 4.7: 58.0 (#18), Inkling: 34.5 (#234)

Coding benchmarks
BenchmarkGrok 4.7Inkling
FrontierCode47.6%14%
LMArena WebDev16391413
FrontierSWE29.5%4.1%
SciCode57.8%47%
LMArena Coding14271464
CursorBench46.3%—
WeirdML—32.3%
ALE-Bench—946

Agentic & Tool Use Grok 4.7 leads

Grok 4.7: 36.7 (#37), Inkling: 29.6 (#85)

Agentic & Tool Use benchmarks
BenchmarkGrok 4.7Inkling
APEX-Agents54.6%33.8%
τ²-bench Banking—25%
GDP.pdf22.8%—
Vending-Bench 210,537—

Reasoning Grok 4.7 leads

Grok 4.7: 49.1 (#40), Inkling: 40.4 (#56)

Reasoning benchmarks
BenchmarkGrok 4.7Inkling
CritPt18%5.4%
Chess Puzzles38%21%
LMArena Hard Prompts14131451
DTBench96%87.5%
LMCA49.4%37.6%
Epoch Capabilities Index153.53148.54
ARC-AGI-2—36.5%
SimpleBench—50%
NYT Connections (extended)76.8%—
ARC-AGI-1—79.5%
Mystery Game Puzzles29%—

Math Grok 4.7 leads

Grok 4.7: 57.8 (#39), Inkling: 31.3 (#225)

Math benchmarks
BenchmarkGrok 4.7Inkling
FrontierMath (Tiers 1-3)53%33.3%
FrontierMath Tier 417.1%4.9%
OTIS Mock AIME 2024-202598.1%88.9%
ProofBench34%0%
LMArena Math14071479

Knowledge Grok 4.7 leads

Grok 4.7: 62.8 (#22), Inkling: 55.1 (#49)

Knowledge benchmarks
BenchmarkGrok 4.7Inkling
GPQA Diamond92.7%88.3%
SimpleQA Verified56%40.3%
LMArena Expert14221465

Multimodal Not comparable

Grok 4.7: 35.5 (#87), Inkling: —

Multimodal benchmarks
BenchmarkGrok 4.7Inkling
LMArena Vision1228—
Blueprint-Bench 232.5%—
Furniture Assembly20.8%—

Multilingual Inkling leads

Grok 4.7: 50.8 (#116), Inkling: 54.0 (#52)

Multilingual benchmarks
BenchmarkGrok 4.7Inkling
LMArena Non-English13891434
LMArena Chinese14551490
LMArena French14551458
LMArena Russian13971429
LMArena Spanish14001448
LMArena German—1446
LMArena Japanese—1429
LMArena Korean—1404

Instruction Following Inkling leads

Grok 4.7: 74.1 (#105), Inkling: 75.1 (#71)

Instruction Following benchmarks
BenchmarkGrok 4.7Inkling
LMArena Instruction Following14041426

Long Context Too close to call

Grok 4.7: 43.1 (#104), Inkling: 43.8 (#86)

Long Context benchmarks
BenchmarkGrok 4.7Inkling
LMArena Longer Query14131434

Writing & Preference Grok 4.7 leads

Grok 4.7: 70.0 (#24), Inkling: 65.2 (#51)

Writing & Preference benchmarks
BenchmarkGrok 4.7Inkling
LMArena Text13991441
LMArena Creative Writing13911387
EQ-Bench Creative Writing20071611
LMArena Multi-Turn13931436
EQ-Bench 4—1226

Frequently asked questions

Is Grok 4.7 better than Inkling?

Grok 4.7 is the stronger model overall, scoring 53.1 to 44.1 on the Noometry Index.

Which is cheaper, Grok 4.7 or Inkling?

Inkling is cheaper. It lists at $1.87 per million input tokens and $4.68 per million output tokens; Grok 4.7 lists at $2 and $6.

Is Grok 4.7 or Inkling better for coding?

Grok 4.7 scores higher on coding benchmarks: 58.0 versus 34.5 in the Noometry coding category.

Which has the bigger context window?

Grok 4.7 does, with 500K tokens against 66K.

How many benchmarks do Grok 4.7 and Inkling share?

31 benchmarks have published results for both models. Grok 4.7 has 39 scored results on Noometry and Inkling has 41.

Related comparisons

Go deeper