Model comparison

DeepSeek-V3 vs Inkling

Inkling is the stronger model overall, scoring 44.1 to 39.5 on the Noometry Index. DeepSeek-V3 costs 6.4× less per token, which makes it the better buy when Inkling's lead doesn't matter for your workload.

Last verified . 27 shared benchmarks.

DeepSeek-V3 DeepSeek

39.5

Rank #166 Confirmed

Inkling Thinking Machines Lab

44.1

Rank #80 Confirmed

Summary

  • They share 27 benchmarks with published results for both. DeepSeek-V3 scores higher in 2 categories and Inkling in 6 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Inkling leads 40.4 to 20.5.
  • The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 37.8% for DeepSeek-V3 and 88.9% for Inkling.
  • DeepSeek-V3 is cheaper at $0.24 / $0.90 per million input/output tokens, against $1.87 / $4.68 for Inkling.
  • DeepSeek-V3 accepts more context: 164K tokens versus 66K.

Side by side

DeepSeek-V3 and Inkling specifications
DeepSeek-V3Inkling
ProviderDeepSeekThinking Machines Lab
Noometry Index39.544.1
Released2024-12-262026-07-15
WeightsOpenOpen
Context window164K66K
Max output164K66K
Input $ / M tokens$0.24$1.87
Output $ / M tokens$0.90$4.68
Results tracked6041

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding DeepSeek-V3 leads

DeepSeek-V3: 42.3 (#106), Inkling: 34.5 (#234)

Coding benchmarks
BenchmarkDeepSeek-V3Inkling
SciCode35.8%47%
WeirdML36.1%32.3%
LMArena Coding13681464
FrontierCode—14%
Aider Polyglot55.1%—
LMArena WebDev—1413
FrontierSWE—4.1%
BigCodeBench Instruct50%—
LiveBench Coding70.9%—
BigCodeBench Complete62.2%—
ALE-Bench—946
HumanEval+86.6%—
MBPP+73%—

Agentic & Tool Use Not comparable

DeepSeek-V3: —, Inkling: 29.6 (#85)

Agentic & Tool Use benchmarks
BenchmarkDeepSeek-V3Inkling
APEX-Agents—33.8%
τ²-bench Banking—25%
METR Time Horizons49.6%—

Reasoning Inkling leads

DeepSeek-V3: 20.5 (#236), Inkling: 40.4 (#56)

Reasoning benchmarks
BenchmarkDeepSeek-V3Inkling
SimpleBench27.2%50%
CritPt0%5.4%
LMArena Hard Prompts13651451
DTBench64.8%87.5%
LMCA15.5%37.6%
Epoch Capabilities Index135.94148.54
ARC-AGI-2—36.5%
Kagi LLM Benchmark52.3%—
ARC-AGI-1—79.5%
Chess Puzzles—21%
LiveBench Reasoning65.8%—
LiveBench Data Analysis60.9%—
BIG-Bench Hard87.5%—
ForecastBench59.1—
HellaSwag88.9%—
LiveBench66.9%—
PIQA84.7%—
WinoGrande85.2%—

Math Too close to call

DeepSeek-V3: 32.1 (#219), Inkling: 31.3 (#225)

Math benchmarks
BenchmarkDeepSeek-V3Inkling
OTIS Mock AIME 2024-202537.8%88.9%
LMArena Math13731479
FrontierMath (Tiers 1-3)—33.3%
FrontierMath Tier 4—4.9%
ProofBench—0%
Omni-MATH40.3%—
LiveBench Math73.5%—
MATH Level 575.5%—
FrontierMath (Feb 2025 set)1.7%—

Knowledge Inkling leads

DeepSeek-V3: 37.5 (#155), Inkling: 55.1 (#49)

Knowledge benchmarks
BenchmarkDeepSeek-V3Inkling
GPQA Diamond67.6%88.3%
LMArena Expert13511465
SimpleQA Verified—40.3%
MMLU-Pro72.3%—
Confabulations26.1%—
Vectara Hallucination Rate6.1%—
GPQA (HELM)53.8%—
ARC (AI2) Challenge95.3%—
MMLU87.2%—
TriviaQA82.9%—

Multilingual Inkling leads

DeepSeek-V3: 48.5 (#143), Inkling: 54.0 (#52)

Multilingual benchmarks
BenchmarkDeepSeek-V3Inkling
LMArena Non-English13581434
LMArena Chinese13911490
LMArena French13851458
LMArena German13741446
LMArena Japanese13331429
LMArena Korean13191404
LMArena Russian13731429
LMArena Spanish13581448

Instruction Following Inkling leads

DeepSeek-V3: 72.8 (#130), Inkling: 75.1 (#71)

Instruction Following benchmarks
BenchmarkDeepSeek-V3Inkling
LMArena Instruction Following13451426
LiveBench Instruction Following81.5%—
IFEval83.2%—

Long Context Inkling leads

DeepSeek-V3: 34.0 (#253), Inkling: 43.8 (#86)

Long Context benchmarks
BenchmarkDeepSeek-V3Inkling
LMArena Longer Query13521434
Fiction.LiveBench50%—

Writing & Preference Inkling leads

DeepSeek-V3: 57.4 (#130), Inkling: 65.2 (#51)

Writing & Preference benchmarks
BenchmarkDeepSeek-V3Inkling
LMArena Text13751441
LMArena Creative Writing13641387
EQ-Bench Creative Writing14721611
LMArena Multi-Turn13891436
Short-Story Creative Writing77%—
WildBench83%—
EQ-Bench 4—1226
LiveBench Language49.1%—

Frequently asked questions

Is DeepSeek-V3 better than Inkling?

Inkling is the stronger model overall, scoring 44.1 to 39.5 on the Noometry Index. DeepSeek-V3 costs 6.4× less per token, which makes it the better buy when Inkling's lead doesn't matter for your workload.

Which is cheaper, DeepSeek-V3 or Inkling?

DeepSeek-V3 is cheaper. It lists at $0.24 per million input tokens and $0.90 per million output tokens; Inkling lists at $1.87 and $4.68.

Is DeepSeek-V3 or Inkling better for coding?

DeepSeek-V3 scores higher on coding benchmarks: 42.3 versus 34.5 in the Noometry coding category.

Which has the bigger context window?

DeepSeek-V3 does, with 164K tokens against 66K.

How many benchmarks do DeepSeek-V3 and Inkling share?

27 benchmarks have published results for both models. DeepSeek-V3 has 60 scored results on Noometry and Inkling has 41.

Related comparisons

Go deeper