Model comparison

Grok 4.5 vs Inkling

Grok 4.5 is the stronger model overall, scoring 55.0 to 44.1 on the Noometry Index.

Last verified . 39 shared benchmarks.

Grok 4.5 xAI

55.0

Rank #25 Confirmed

Inkling Thinking Machines Lab

44.1

Rank #80 Confirmed

Summary

  • They share 39 benchmarks with published results for both. Grok 4.5 scores higher in 9 categories and Inkling in 0 categories; 5 gaps are clear of the uncertainty.
  • The widest gap is in math, where Grok 4.5 leads 60.9 to 31.3.
  • The biggest single-benchmark swing is ProofBench: 31% for Grok 4.5 and 0% for Inkling.
  • Inkling is cheaper at $1.87 / $4.68 per million input/output tokens, against $2 / $6 for Grok 4.5.
  • Grok 4.5 accepts more context: 500K tokens versus 66K.
  • Inkling has downloadable open weights; the other is API-only.

Side by side

Grok 4.5 and Inkling specifications
Grok 4.5Inkling
ProviderxAIThinking Machines Lab
Noometry Index55.044.1
Released2026-07-082026-07-15
WeightsProprietaryOpen
Context window500K66K
Max output500K66K
Input $ / M tokens$2$1.87
Output $ / M tokens$6$4.68
Results tracked5241

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Grok 4.5 leads

Grok 4.5: 52.2 (#35), Inkling: 34.5 (#234)

Coding benchmarks
BenchmarkGrok 4.5Inkling
FrontierCode42.4%14%
LMArena WebDev15531413
SciCode54.1%47%
WeirdML46.4%32.3%
LMArena Coding14741464
ALE-Bench1,309946
DeepSWE53.8%—
FrontierSWE—4.1%

Agentic & Tool Use Grok 4.5 leads

Grok 4.5: 44.4 (#17), Inkling: 29.6 (#85)

Agentic & Tool Use benchmarks
BenchmarkGrok 4.5Inkling
APEX-Agents56.2%33.8%
τ²-bench Banking47.9%25%
PostTrainBench23.4%—
GBAEval65.4%—
GDP.pdf14%—
LMArena Search1213—
Vending-Bench 23,887—

Reasoning Grok 4.5 leads

Grok 4.5: 56.1 (#25), Inkling: 40.4 (#56)

Reasoning benchmarks
BenchmarkGrok 4.5Inkling
ARC-AGI-252.6%36.5%
SimpleBench70%50%
ARC-AGI-187.2%79.5%
CritPt15.4%5.4%
Chess Puzzles36%21%
LMArena Hard Prompts14621451
DTBench96.5%87.5%
LMCA45.2%37.6%
Epoch Capabilities Index153.92148.54
Kagi LLM Benchmark83.5%—
NYT Connections (extended)79.9%—
Surface Evolver Bench74.4%—

Math Grok 4.5 leads

Grok 4.5: 60.9 (#35), Inkling: 31.3 (#225)

Math benchmarks
BenchmarkGrok 4.5Inkling
FrontierMath (Tiers 1-3)57.2%33.3%
FrontierMath Tier 424.4%4.9%
OTIS Mock AIME 2024-202597.8%88.9%
ProofBench31%0%
LMArena Math14591479

Knowledge Grok 4.5 leads

Grok 4.5: 62.3 (#24), Inkling: 55.1 (#49)

Knowledge benchmarks
BenchmarkGrok 4.5Inkling
GPQA Diamond93.4%88.3%
SimpleQA Verified48.3%40.3%
LMArena Expert14661465

Multimodal Not comparable

Grok 4.5: 37.6 (#72), Inkling: —

Multimodal benchmarks
BenchmarkGrok 4.5Inkling
LMArena Vision1288—
Blueprint-Bench 227.3%—
Furniture Assembly22.5%—
LMArena Document1452—

Multilingual Too close to call

Grok 4.5: 54.4 (#42), Inkling: 54.0 (#52)

Multilingual benchmarks
BenchmarkGrok 4.5Inkling
LMArena Non-English14401434
LMArena Chinese14961490
LMArena French14561458
LMArena German14461446
LMArena Japanese14281429
LMArena Korean14041404
LMArena Russian14481429
LMArena Spanish14501448

Instruction Following Too close to call

Grok 4.5: 76.0 (#48), Inkling: 75.1 (#71)

Instruction Following benchmarks
BenchmarkGrok 4.5Inkling
LMArena Instruction Following14461426

Long Context Too close to call

Grok 4.5: 44.8 (#56), Inkling: 43.8 (#86)

Long Context benchmarks
BenchmarkGrok 4.5Inkling
LMArena Longer Query14631434

Writing & Preference Too close to call

Grok 4.5: 65.8 (#42), Inkling: 65.2 (#51)

Writing & Preference benchmarks
BenchmarkGrok 4.5Inkling
LMArena Text14481441
LMArena Creative Writing14421387
EQ-Bench Creative Writing15791611
LMArena Multi-Turn14561436
EQ-Bench 4—1226

Frequently asked questions

Is Grok 4.5 better than Inkling?

Grok 4.5 is the stronger model overall, scoring 55.0 to 44.1 on the Noometry Index.

Which is cheaper, Grok 4.5 or Inkling?

Inkling is cheaper. It lists at $1.87 per million input tokens and $4.68 per million output tokens; Grok 4.5 lists at $2 and $6.

Is Grok 4.5 or Inkling better for coding?

Grok 4.5 scores higher on coding benchmarks: 52.2 versus 34.5 in the Noometry coding category.

Which has the bigger context window?

Grok 4.5 does, with 500K tokens against 66K.

How many benchmarks do Grok 4.5 and Inkling share?

39 benchmarks have published results for both models. Grok 4.5 has 52 scored results on Noometry and Inkling has 41.

Related comparisons

Go deeper