Model comparison

Grok 4.3 vs Inkling

Grok 4.3 and Inkling score almost the same on the Noometry Index (43.8 vs 44.1), so choose on price, context window or the category you care about most.

Last verified . 33 shared benchmarks.

Grok 4.3 xAI

43.8

Rank #86 Confirmed

Inkling Thinking Machines Lab

44.1

Rank #80 Confirmed

Summary

  • They share 33 benchmarks with published results for both. Grok 4.3 scores higher in 2 categories and Inkling in 7 categories; 9 gaps are clear of the uncertainty.
  • The widest gap is in math, where Grok 4.3 leads 46.0 to 31.3.
  • The biggest single-benchmark swing is WeirdML: 49.9% for Grok 4.3 and 32.3% for Inkling.
  • Grok 4.3 is cheaper at $1.25 / $2.50 per million input/output tokens, against $1.87 / $4.68 for Inkling.
  • Grok 4.3 accepts more context: 1M tokens versus 66K.
  • Inkling has downloadable open weights; the other is API-only.

Side by side

Grok 4.3 and Inkling specifications
Grok 4.3Inkling
ProviderxAIThinking Machines Lab
Noometry Index43.844.1
Released2026-04-172026-07-15
WeightsProprietaryOpen
Context window1M66K
Max output30K66K
Input $ / M tokens$1.25$1.87
Output $ / M tokens$2.50$4.68
Results tracked4041

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Grok 4.3 leads

Grok 4.3: 41.6 (#121), Inkling: 34.5 (#234)

Coding benchmarks
BenchmarkGrok 4.3Inkling
LMArena WebDev13571413
SciCode47.3%47%
WeirdML49.9%32.3%
LMArena Coding14151464
ALE-Bench944.17946
FrontierCode—14%
FrontierSWE—4.1%

Agentic & Tool Use Inkling leads

Grok 4.3: 27.7 (#99), Inkling: 29.6 (#85)

Agentic & Tool Use benchmarks
BenchmarkGrok 4.3Inkling
APEX-Agents—33.8%
τ²-bench Banking—25%
GDP.pdf8%—
LMArena Search1165—
Vending-Bench 235.26—

Reasoning Inkling leads

Grok 4.3: 35.9 (#68), Inkling: 40.4 (#56)

Reasoning benchmarks
BenchmarkGrok 4.3Inkling
CritPt8%5.4%
Chess Puzzles25%21%
LMArena Hard Prompts13961451
DTBench90.7%87.5%
LMCA38.3%37.6%
Epoch Capabilities Index149.16148.54
ARC-AGI-2—36.5%
SimpleBench—50%
NYT Connections (extended)55.2%—
ARC-AGI-1—79.5%
ForecastBench60.3—

Math Grok 4.3 leads

Grok 4.3: 46.0 (#74), Inkling: 31.3 (#225)

Math benchmarks
BenchmarkGrok 4.3Inkling
FrontierMath (Tiers 1-3)42.8%33.3%
FrontierMath Tier 414.6%4.9%
OTIS Mock AIME 2024-202593.3%88.9%
ProofBench11%0%
LMArena Math13881479

Knowledge Inkling leads

Grok 4.3: 52.5 (#62), Inkling: 55.1 (#49)

Knowledge benchmarks
BenchmarkGrok 4.3Inkling
GPQA Diamond88.8%88.3%
SimpleQA Verified33.2%40.3%
LMArena Expert13851465

Multimodal Not comparable

Grok 4.3: 31.6 (#104), Inkling: —

Multimodal benchmarks
BenchmarkGrok 4.3Inkling
LMArena Vision1229—
Blueprint-Bench 20%—

Multilingual Inkling leads

Grok 4.3: 50.5 (#120), Inkling: 54.0 (#52)

Multilingual benchmarks
BenchmarkGrok 4.3Inkling
LMArena Non-English13851434
LMArena Chinese14221490
LMArena French14121458
LMArena German13951446
LMArena Japanese13791429
LMArena Korean13561404
LMArena Russian13991429
LMArena Spanish13981448

Instruction Following Inkling leads

Grok 4.3: 72.1 (#140), Inkling: 75.1 (#71)

Instruction Following benchmarks
BenchmarkGrok 4.3Inkling
LMArena Instruction Following13661426

Long Context Inkling leads

Grok 4.3: 42.5 (#123), Inkling: 43.8 (#86)

Long Context benchmarks
BenchmarkGrok 4.3Inkling
LMArena Longer Query13931434

Writing & Preference Inkling leads

Grok 4.3: 58.5 (#118), Inkling: 65.2 (#51)

Writing & Preference benchmarks
BenchmarkGrok 4.3Inkling
LMArena Text13971441
LMArena Creative Writing13801387
EQ-Bench 410751226
LMArena Multi-Turn14061436
EQ-Bench Creative Writing—1611

Frequently asked questions

Is Grok 4.3 better than Inkling?

Grok 4.3 and Inkling score almost the same on the Noometry Index (43.8 vs 44.1), so choose on price, context window or the category you care about most.

Which is cheaper, Grok 4.3 or Inkling?

Grok 4.3 is cheaper. It lists at $1.25 per million input tokens and $2.50 per million output tokens; Inkling lists at $1.87 and $4.68.

Is Grok 4.3 or Inkling better for coding?

Grok 4.3 scores higher on coding benchmarks: 41.6 versus 34.5 in the Noometry coding category.

Which has the bigger context window?

Grok 4.3 does, with 1M tokens against 66K.

How many benchmarks do Grok 4.3 and Inkling share?

33 benchmarks have published results for both models. Grok 4.3 has 40 scored results on Noometry and Inkling has 41.

Related comparisons

Go deeper