Model comparison

DeepSeek-V3.1-Terminus vs Inkling

DeepSeek-V3.1-Terminus and Inkling score almost the same on the Noometry Index (43.1 vs 44.1), so choose on price, context window or the category you care about most.

Last verified . 15 shared benchmarks.

DeepSeek-V3.1-Terminus DeepSeek

43.1

Rank #97 Confirmed

Inkling Thinking Machines Lab

44.1

Rank #80 Confirmed

Summary

  • They share 15 benchmarks with published results for both. DeepSeek-V3.1-Terminus scores higher in 2 categories and Inkling in 5 categories; 6 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Inkling leads 40.4 to 26.4.
  • The biggest single-benchmark swing is LMCA: 28.6% for DeepSeek-V3.1-Terminus and 37.6% for Inkling.
  • DeepSeek-V3.1-Terminus is cheaper at $0.27 / $1 per million input/output tokens, against $1.87 / $4.68 for Inkling.
  • DeepSeek-V3.1-Terminus accepts more context: 164K tokens versus 66K.

Side by side

DeepSeek-V3.1-Terminus and Inkling specifications
DeepSeek-V3.1-TerminusInkling
ProviderDeepSeekThinking Machines Lab
Noometry Index43.144.1
Released2025-09-222026-07-15
WeightsOpenOpen
Context window164K66K
Max output147K66K
Input $ / M tokens$0.27$1.87
Output $ / M tokens$1$4.68
Results tracked1641

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding DeepSeek-V3.1-Terminus leads

DeepSeek-V3.1-Terminus: 42.0 (#113), Inkling: 34.5 (#234)

Coding benchmarks
BenchmarkDeepSeek-V3.1-TerminusInkling
SciCode40.6%47%
LMArena Coding14261464
ALE-Bench745.17946
FrontierCode—14%
LMArena WebDev—1413
FrontierSWE—4.1%
WeirdML—32.3%

Agentic & Tool Use Not comparable

DeepSeek-V3.1-Terminus: —, Inkling: 29.6 (#85)

Agentic & Tool Use benchmarks
BenchmarkDeepSeek-V3.1-TerminusInkling
APEX-Agents—33.8%
τ²-bench Banking—25%

Reasoning Inkling leads

DeepSeek-V3.1-Terminus: 26.4 (#133), Inkling: 40.4 (#56)

Reasoning benchmarks
BenchmarkDeepSeek-V3.1-TerminusInkling
CritPt1.7%5.4%
LMArena Hard Prompts14261451
DTBench81.3%87.5%
LMCA28.6%37.6%
ARC-AGI-2—36.5%
SimpleBench—50%
Kagi LLM Benchmark57.4%—
ARC-AGI-1—79.5%
Chess Puzzles—21%
Epoch Capabilities Index—148.54

Math DeepSeek-V3.1-Terminus leads

DeepSeek-V3.1-Terminus: 38.5 (#137), Inkling: 31.3 (#225)

Math benchmarks
BenchmarkDeepSeek-V3.1-TerminusInkling
LMArena Math14021479
FrontierMath (Tiers 1-3)—33.3%
FrontierMath Tier 4—4.9%
OTIS Mock AIME 2024-2025—88.9%
ProofBench—0%

Knowledge Not comparable

DeepSeek-V3.1-Terminus: —, Inkling: 55.1 (#49)

Knowledge benchmarks
BenchmarkDeepSeek-V3.1-TerminusInkling
GPQA Diamond—88.3%
SimpleQA Verified—40.3%
LMArena Expert—1465

Multilingual Inkling leads

DeepSeek-V3.1-Terminus: 52.1 (#92), Inkling: 54.0 (#52)

Multilingual benchmarks
BenchmarkDeepSeek-V3.1-TerminusInkling
LMArena Non-English14071434
LMArena Russian14361429
LMArena Chinese—1490
LMArena French—1458
LMArena German—1446
LMArena Japanese—1429
LMArena Korean—1404
LMArena Spanish—1448

Instruction Following Inkling leads

DeepSeek-V3.1-Terminus: 74.0 (#106), Inkling: 75.1 (#71)

Instruction Following benchmarks
BenchmarkDeepSeek-V3.1-TerminusInkling
LMArena Instruction Following14041426

Long Context Too close to call

DeepSeek-V3.1-Terminus: 43.4 (#97), Inkling: 43.8 (#86)

Long Context benchmarks
BenchmarkDeepSeek-V3.1-TerminusInkling
LMArena Longer Query14211434

Writing & Preference Inkling leads

DeepSeek-V3.1-Terminus: 61.0 (#92), Inkling: 65.2 (#51)

Writing & Preference benchmarks
BenchmarkDeepSeek-V3.1-TerminusInkling
LMArena Text14191441
LMArena Creative Writing14031387
LMArena Multi-Turn14111436
EQ-Bench Creative Writing—1611
EQ-Bench 4—1226

Frequently asked questions

Is DeepSeek-V3.1-Terminus better than Inkling?

DeepSeek-V3.1-Terminus and Inkling score almost the same on the Noometry Index (43.1 vs 44.1), so choose on price, context window or the category you care about most.

Which is cheaper, DeepSeek-V3.1-Terminus or Inkling?

DeepSeek-V3.1-Terminus is cheaper. It lists at $0.27 per million input tokens and $1 per million output tokens; Inkling lists at $1.87 and $4.68.

Is DeepSeek-V3.1-Terminus or Inkling better for coding?

DeepSeek-V3.1-Terminus scores higher on coding benchmarks: 42.0 versus 34.5 in the Noometry coding category.

Which has the bigger context window?

DeepSeek-V3.1-Terminus does, with 164K tokens against 66K.

How many benchmarks do DeepSeek-V3.1-Terminus and Inkling share?

15 benchmarks have published results for both models. DeepSeek-V3.1-Terminus has 16 scored results on Noometry and Inkling has 41.

Related comparisons

Go deeper