Model comparison

Inkling vs o1

Inkling is the stronger model overall, scoring 44.1 to 40.9 on the Noometry Index.

Last verified . 28 shared benchmarks.

Inkling Thinking Machines Lab

44.1

Rank #80 Confirmed

o1 OpenAI

40.9

Rank #143 Confirmed

Summary

  • They share 28 benchmarks with published results for both. Inkling scores higher in 6 categories and o1 in 3 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Inkling leads 55.1 to 41.5.
  • The biggest single-benchmark swing is ARC-AGI-1: 79.5% for Inkling and 30.7% for o1.
  • Inkling is cheaper at $1.87 / $4.68 per million input/output tokens, against $15 / $60 for o1.
  • o1 accepts more context: 200K tokens versus 66K.
  • Inkling has downloadable open weights; the other is API-only.

Side by side

Inkling and o1 specifications
Inklingo1
ProviderThinking Machines LabOpenAI
Noometry Index44.140.9
Released2026-07-152024-09-12
WeightsOpenProprietary
Context window66K200K
Max output66K100K
Input $ / M tokens$1.87$15
Output $ / M tokens$4.68$60
Results tracked4152

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding o1 leads

Inkling: 34.5 (#234), o1: 46.1 (#70)

Coding benchmarks
BenchmarkInklingo1
WeirdML32.3%47.6%
LMArena Coding14641367
FrontierCode14%—
Aider Polyglot—61.7%
LMArena WebDev1413—
FrontierSWE4.1%—
SciCode47%—
LiveBench Coding—69.7%
CadEval—56%
ALE-Bench946—
HumanEval+—89%
MBPP+—80.2%

Agentic & Tool Use Inkling leads

Inkling: 29.6 (#85), o1: 24.6 (#117)

Agentic & Tool Use benchmarks
BenchmarkInklingo1
APEX-Agents33.8%—
τ²-bench Banking25%—
Cybench—10%
METR Time Horizons—51.1%

Reasoning Inkling leads

Inkling: 40.4 (#56), o1: 27.9 (#111)

Reasoning benchmarks
BenchmarkInklingo1
SimpleBench50%41.7%
ARC-AGI-179.5%30.7%
Chess Puzzles21%15%
LMArena Hard Prompts14511371
DTBench87.5%74.7%
LMCA37.6%22.3%
Epoch Capabilities Index148.54141.91
ARC-AGI-236.5%—
CritPt5.4%—
EnigmaEval—5.7%
LiveBench Reasoning—91.6%
LiveBench Data Analysis—65.5%
LiveBench—75.7%

Math o1 leads

Inkling: 31.3 (#225), o1: 36.1 (#175)

Math benchmarks
BenchmarkInklingo1
FrontierMath (Tiers 1-3)33.3%14.7%
OTIS Mock AIME 2024-202588.9%73.3%
LMArena Math14791388
FrontierMath Tier 44.9%—
ProofBench0%—
LiveBench Math—80.3%
MATH Level 5—94.7%
FrontierMath (Feb 2025 set)—9.3%

Knowledge Inkling leads

Inkling: 55.1 (#49), o1: 41.5 (#110)

Knowledge benchmarks
BenchmarkInklingo1
GPQA Diamond88.3%76.8%
SimpleQA Verified40.3%41.1%
LMArena Expert14651361
Humanity's Last Exam—8%
Confabulations—11.7%

Multimodal Not comparable

Inkling: —, o1: 34.2 (#93)

Multimodal benchmarks
BenchmarkInklingo1
LMArena Vision—1168
GeoBench—80%
VPCT—37%
SpatialViz-Bench—41.4%

Multilingual Inkling leads

Inkling: 54.0 (#52), o1: 48.6 (#142)

Multilingual benchmarks
BenchmarkInklingo1
LMArena Non-English14341358
LMArena Chinese14901394
LMArena French14581344
LMArena German14461337
LMArena Japanese14291346
LMArena Korean14041396
LMArena Russian14291356
LMArena Spanish14481345

Instruction Following Too close to call

Inkling: 75.1 (#71), o1: 74.8 (#86)

Instruction Following benchmarks
BenchmarkInklingo1
LMArena Instruction Following14261367
LiveBench Instruction Following—81.5%

Long Context o1 leads

Inkling: 43.8 (#86), o1: 50.3 (#9)

Long Context benchmarks
BenchmarkInklingo1
LMArena Longer Query14341378
Fiction.LiveBench—83.3%

Writing & Preference Inkling leads

Inkling: 65.2 (#51), o1: 55.6 (#144)

Writing & Preference benchmarks
BenchmarkInklingo1
LMArena Text14411366
LMArena Creative Writing13871348
LMArena Multi-Turn14361369
Short-Story Creative Writing—70.2%
EQ-Bench Creative Writing1611—
EQ-Bench 41226—
LiveBench Language—65.4%

Frequently asked questions

Is Inkling better than o1?

Inkling is the stronger model overall, scoring 44.1 to 40.9 on the Noometry Index.

Which is cheaper, Inkling or o1?

Inkling is cheaper. It lists at $1.87 per million input tokens and $4.68 per million output tokens; o1 lists at $15 and $60.

Is Inkling or o1 better for coding?

o1 scores higher on coding benchmarks: 46.1 versus 34.5 in the Noometry coding category.

Which has the bigger context window?

o1 does, with 200K tokens against 66K.

How many benchmarks do Inkling and o1 share?

28 benchmarks have published results for both models. Inkling has 41 scored results on Noometry and o1 has 52.

Related comparisons

Go deeper