Model comparison

DeepSeek-V3.2-Exp vs Inkling

DeepSeek-V3.2-Exp and Inkling score almost the same on the Noometry Index (44.3 vs 44.1), so choose on price, context window or the category you care about most.

Last verified . 32 shared benchmarks.

DeepSeek-V3.2-Exp DeepSeek

44.3

Rank #78 Confirmed

Inkling Thinking Machines Lab

44.1

Rank #80 Confirmed

Summary

  • They share 32 benchmarks with published results for both. DeepSeek-V3.2-Exp scores higher in 4 categories and Inkling in 5 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Inkling leads 40.4 to 22.1.
  • The biggest single-benchmark swing is ARC-AGI-2: 4% for DeepSeek-V3.2-Exp and 36.5% for Inkling.
  • DeepSeek-V3.2-Exp is cheaper at $0.26 / $0.38 per million input/output tokens, against $1.87 / $4.68 for Inkling.
  • DeepSeek-V3.2-Exp accepts more context: 164K tokens versus 66K.

Side by side

DeepSeek-V3.2-Exp and Inkling specifications
DeepSeek-V3.2-ExpInkling
ProviderDeepSeekThinking Machines Lab
Noometry Index44.344.1
Released2025-09-292026-07-15
WeightsOpenOpen
Context window164K66K
Max output66K66K
Input $ / M tokens$0.26$1.87
Output $ / M tokens$0.38$4.68
Results tracked4941

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding DeepSeek-V3.2-Exp leads

DeepSeek-V3.2-Exp: 46.5 (#65), Inkling: 34.5 (#234)

Coding benchmarks
BenchmarkDeepSeek-V3.2-ExpInkling
LMArena WebDev13621413
SciCode38.9%47%
WeirdML39.5%32.3%
LMArena Coding14541464
FrontierCode—14%
SWE-bench Verified (bash only)70%—
Aider Polyglot74.2%—
SWE-bench Multilingual59%—
FrontierSWE—4.1%
ALE-Bench—946

Agentic & Tool Use DeepSeek-V3.2-Exp leads

DeepSeek-V3.2-Exp: 32.7 (#59), Inkling: 29.6 (#85)

Agentic & Tool Use benchmarks
BenchmarkDeepSeek-V3.2-ExpInkling
APEX-Agents21.3%33.8%
Terminal-Bench39.6%—
Berkeley Function Calling Leaderboard56.7%—
TheAgentCompany42.9%—
τ²-bench Banking—25%
Vending-Bench 21,034—

Reasoning Inkling leads

DeepSeek-V3.2-Exp: 22.1 (#208), Inkling: 40.4 (#56)

Reasoning benchmarks
BenchmarkDeepSeek-V3.2-ExpInkling
ARC-AGI-24%36.5%
ARC-AGI-157%79.5%
CritPt2.9%5.4%
Chess Puzzles14%21%
LMArena Hard Prompts14341451
DTBench87.7%87.5%
LMCA29.1%37.6%
Epoch Capabilities Index146.27148.54
SimpleBench—50%
Kagi LLM Benchmark52.2%—
NYT Connections (extended)36.7%—
Thematic Generalization65%—

Math DeepSeek-V3.2-Exp leads

DeepSeek-V3.2-Exp: 41.7 (#87), Inkling: 31.3 (#225)

Knowledge Inkling leads

DeepSeek-V3.2-Exp: 51.7 (#66), Inkling: 55.1 (#49)

Knowledge benchmarks
BenchmarkDeepSeek-V3.2-ExpInkling
GPQA Diamond83.4%88.3%
LMArena Expert14361465
SimpleQA Verified—40.3%
Vectara Hallucination Rate5.3%—

Multilingual Inkling leads

DeepSeek-V3.2-Exp: 52.2 (#90), Inkling: 54.0 (#52)

Multilingual benchmarks
BenchmarkDeepSeek-V3.2-ExpInkling
LMArena Non-English14091434
LMArena Chinese14611490
LMArena French14331458
LMArena German14401446
LMArena Japanese13741429
LMArena Korean13711404
LMArena Russian14241429
LMArena Spanish14401448

Instruction Following Too close to call

DeepSeek-V3.2-Exp: 74.5 (#93), Inkling: 75.1 (#71)

Instruction Following benchmarks
BenchmarkDeepSeek-V3.2-ExpInkling
LMArena Instruction Following14131426

Long Context DeepSeek-V3.2-Exp leads

DeepSeek-V3.2-Exp: 47.6 (#16), Inkling: 43.8 (#86)

Long Context benchmarks
BenchmarkDeepSeek-V3.2-ExpInkling
LMArena Longer Query14281434
Fiction.LiveBench83.3%—
CL-bench13.2%—
CL-bench Life9.5%—

Writing & Preference Inkling leads

DeepSeek-V3.2-Exp: 62.4 (#77), Inkling: 65.2 (#51)

Writing & Preference benchmarks
BenchmarkDeepSeek-V3.2-ExpInkling
LMArena Text14251441
LMArena Creative Writing14031387
EQ-Bench Creative Writing15151611
LMArena Multi-Turn14271436
EQ-Bench 4—1226

Frequently asked questions

Is DeepSeek-V3.2-Exp better than Inkling?

DeepSeek-V3.2-Exp and Inkling score almost the same on the Noometry Index (44.3 vs 44.1), so choose on price, context window or the category you care about most.

Which is cheaper, DeepSeek-V3.2-Exp or Inkling?

DeepSeek-V3.2-Exp is cheaper. It lists at $0.26 per million input tokens and $0.38 per million output tokens; Inkling lists at $1.87 and $4.68.

Is DeepSeek-V3.2-Exp or Inkling better for coding?

DeepSeek-V3.2-Exp scores higher on coding benchmarks: 46.5 versus 34.5 in the Noometry coding category.

Which has the bigger context window?

DeepSeek-V3.2-Exp does, with 164K tokens against 66K.

How many benchmarks do DeepSeek-V3.2-Exp and Inkling share?

32 benchmarks have published results for both models. DeepSeek-V3.2-Exp has 49 scored results on Noometry and Inkling has 41.

Related comparisons

Go deeper