Model comparison

GPT-5.4 mini vs Inkling

GPT-5.4 mini and Inkling score almost the same on the Noometry Index (45.0 vs 44.1), so choose on price, context window or the category you care about most.

Last verified . 36 shared benchmarks.

GPT-5.4 mini OpenAI

45.0

Rank #76 Confirmed

Inkling Thinking Machines Lab

44.1

Rank #80 Confirmed

Summary

  • They share 36 benchmarks with published results for both. GPT-5.4 mini scores higher in 3 categories and Inkling in 6 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in math, where GPT-5.4 mini leads 45.5 to 31.3.
  • The biggest single-benchmark swing is WeirdML: 60.3% for GPT-5.4 mini and 32.3% for Inkling.
  • GPT-5.4 mini is cheaper at $0.75 / $4.50 per million input/output tokens, against $1.87 / $4.68 for Inkling.
  • GPT-5.4 mini accepts more context: 400K tokens versus 66K.
  • Inkling has downloadable open weights; the other is API-only.

Side by side

GPT-5.4 mini and Inkling specifications
GPT-5.4 miniInkling
ProviderOpenAIThinking Machines Lab
Noometry Index45.044.1
Released2026-03-172026-07-15
WeightsProprietaryOpen
Context window400K66K
Max output128K66K
Input $ / M tokens$0.75$1.87
Output $ / M tokens$4.50$4.68
Results tracked4641

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding GPT-5.4 mini leads

GPT-5.4 mini: 45.2 (#72), Inkling: 34.5 (#234)

Coding benchmarks
BenchmarkGPT-5.4 miniInkling
FrontierCode27%14%
LMArena WebDev13971413
SciCode49.9%47%
WeirdML60.3%32.3%
LMArena Coding14381464
ALE-Bench1,189946
FrontierSWE—4.1%

Agentic & Tool Use Too close to call

GPT-5.4 mini: 29.9 (#81), Inkling: 29.6 (#85)

Agentic & Tool Use benchmarks
BenchmarkGPT-5.4 miniInkling
APEX-Agents—33.8%
τ²-bench Banking—25%
DeepResearch Bench36.3%—

Reasoning Inkling leads

GPT-5.4 mini: 30.4 (#85), Inkling: 40.4 (#56)

Reasoning benchmarks
BenchmarkGPT-5.4 miniInkling
ARC-AGI-218.9%36.5%
ARC-AGI-163.7%79.5%
CritPt10%5.4%
Chess Puzzles24%21%
LMArena Hard Prompts14241451
DTBench80%87.5%
LMCA40.8%37.6%
Epoch Capabilities Index148.84148.54
SimpleBench—50%
Kagi LLM Benchmark37.9%—
NYT Connections (extended)61.8%—
Thematic Generalization61.7%—
Mystery Game Puzzles11%—
ForecastBench57—

Math GPT-5.4 mini leads

GPT-5.4 mini: 45.5 (#75), Inkling: 31.3 (#225)

Math benchmarks
BenchmarkGPT-5.4 miniInkling
FrontierMath (Tiers 1-3)51.2%33.3%
FrontierMath Tier 49.8%4.9%
OTIS Mock AIME 2024-202588.9%88.9%
ProofBench21%0%
LMArena Math14191479
FrontierMath (Feb 2025 set)28.3%—
FrontierMath Tier 4 (v1)2.1%—

Knowledge Inkling leads

GPT-5.4 mini: 51.5 (#67), Inkling: 55.1 (#49)

Knowledge benchmarks
BenchmarkGPT-5.4 miniInkling
GPQA Diamond86.9%88.3%
SimpleQA Verified29.4%40.3%
LMArena Expert14351465
Vectara Hallucination Rate5.5%—

Multimodal Not comparable

GPT-5.4 mini: 39.7 (#56), Inkling: —

Multimodal benchmarks
BenchmarkGPT-5.4 miniInkling
LMArena Vision1245—

Multilingual Inkling leads

GPT-5.4 mini: 51.9 (#96), Inkling: 54.0 (#52)

Multilingual benchmarks
BenchmarkGPT-5.4 miniInkling
LMArena Non-English14051434
LMArena Chinese14461490
LMArena French14401458
LMArena German14091446
LMArena Japanese13741429
LMArena Korean13681404
LMArena Russian14171429
LMArena Spanish14051448

Instruction Following Inkling leads

GPT-5.4 mini: 74.1 (#102), Inkling: 75.1 (#71)

Instruction Following benchmarks
BenchmarkGPT-5.4 miniInkling
LMArena Instruction Following14051426

Long Context Too close to call

GPT-5.4 mini: 43.0 (#112), Inkling: 43.8 (#86)

Long Context benchmarks
BenchmarkGPT-5.4 miniInkling
LMArena Longer Query14071434

Writing & Preference Inkling leads

GPT-5.4 mini: 64.0 (#58), Inkling: 65.2 (#51)

Writing & Preference benchmarks
BenchmarkGPT-5.4 miniInkling
LMArena Text14121441
LMArena Creative Writing13701387
EQ-Bench Creative Writing16651611
LMArena Multi-Turn14291436
EQ-Bench 4—1226

Frequently asked questions

Is GPT-5.4 mini better than Inkling?

GPT-5.4 mini and Inkling score almost the same on the Noometry Index (45.0 vs 44.1), so choose on price, context window or the category you care about most.

Which is cheaper, GPT-5.4 mini or Inkling?

GPT-5.4 mini is cheaper. It lists at $0.75 per million input tokens and $4.50 per million output tokens; Inkling lists at $1.87 and $4.68.

Is GPT-5.4 mini or Inkling better for coding?

GPT-5.4 mini scores higher on coding benchmarks: 45.2 versus 34.5 in the Noometry coding category.

Which has the bigger context window?

GPT-5.4 mini does, with 400K tokens against 66K.

How many benchmarks do GPT-5.4 mini and Inkling share?

36 benchmarks have published results for both models. GPT-5.4 mini has 46 scored results on Noometry and Inkling has 41.

Related comparisons

Go deeper