Model comparison

Inkling-Small vs Longcat Flash Chat

Inkling-Small is the stronger model overall, scoring 46.5 to 42.1 on the Noometry Index.

Last verified . 17 shared benchmarks.

Inkling-Small Thinking Machines Lab

46.5

Rank #63 Confirmed

Longcat Flash Chat Meituan

42.1

Rank #120 Confirmed

Summary

  • They share 17 benchmarks with published results for both. Inkling-Small scores higher in 4 categories and Longcat Flash Chat in 4 categories; 4 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Inkling-Small leads 38.6 to 19.0.

Side by side

Inkling-Small and Longcat Flash Chat specifications
Inkling-SmallLongcat Flash Chat
ProviderThinking Machines LabMeituan
Noometry Index46.542.1
Released2026-07-15—
WeightsOpenOpen
Context window524K—
Max output1.05M—
Input $ / M tokens$0.45—
Output $ / M tokens$1.20—
Results tracked3319

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

Inkling-Small: 43.6 (#85), Longcat Flash Chat: 43.5 (#87)

Coding benchmarks
BenchmarkInkling-SmallLongcat Flash Chat
LMArena Coding14511471
LMArena WebDev1409—
SciCode48.7%—

Reasoning Inkling-Small leads

Inkling-Small: 38.6 (#63), Longcat Flash Chat: 19.0 (#272)

Reasoning benchmarks
BenchmarkInkling-SmallLongcat Flash Chat
LMArena Hard Prompts14231440
ARC-AGI-240.1%—
Kagi LLM Benchmark—43.9%
NYT Connections (extended)—17.7%
ARC-AGI-184%—
CritPt8.3%—
Chess Puzzles18%—
Mystery Game Puzzles6%—
Epoch Capabilities Index150.15—

Math Inkling-Small leads

Inkling-Small: 45.1 (#77), Longcat Flash Chat: 39.4 (#107)

Math benchmarks
BenchmarkInkling-SmallLongcat Flash Chat
LMArena Math14591442
FrontierMath (Tiers 1-3)46.3%—
FrontierMath Tier 417.1%—
OTIS Mock AIME 2024-202590%—
ProofBench6%—

Knowledge Inkling-Small leads

Inkling-Small: 48.2 (#77), Longcat Flash Chat: 40.6 (#116)

Knowledge benchmarks
BenchmarkInkling-SmallLongcat Flash Chat
LMArena Expert14421454
GPQA Diamond88.5%—
SimpleQA Verified19.1%—

Multimodal Not comparable

Inkling-Small: 39.1 (#62), Longcat Flash Chat: —

Multimodal benchmarks
BenchmarkInkling-SmallLongcat Flash Chat
LMArena Vision1235—

Multilingual Too close to call

Inkling-Small: 51.7 (#104), Longcat Flash Chat: 51.9 (#101)

Multilingual benchmarks
BenchmarkInkling-SmallLongcat Flash Chat
LMArena Non-English14021404
LMArena Chinese14651465
LMArena French14361456
LMArena German14051408
LMArena Japanese14051373
LMArena Korean13631371
LMArena Russian13911395
LMArena Spanish14281445

Instruction Following Too close to call

Inkling-Small: 73.8 (#114), Longcat Flash Chat: 74.4 (#96)

Instruction Following benchmarks
BenchmarkInkling-SmallLongcat Flash Chat
LMArena Instruction Following13991411

Long Context Too close to call

Inkling-Small: 42.7 (#118), Longcat Flash Chat: 43.5 (#93)

Long Context benchmarks
BenchmarkInkling-SmallLongcat Flash Chat
LMArena Longer Query14011425

Writing & Preference Longcat Flash Chat leads

Inkling-Small: 59.6 (#107), Longcat Flash Chat: 61.0 (#91)

Writing & Preference benchmarks
BenchmarkInkling-SmallLongcat Flash Chat
LMArena Text14141427
LMArena Creative Writing13311388
LMArena Multi-Turn14181418
EQ-Bench Creative Writing1491—

Frequently asked questions

Is Inkling-Small better than Longcat Flash Chat?

Inkling-Small is the stronger model overall, scoring 46.5 to 42.1 on the Noometry Index.

Is Inkling-Small or Longcat Flash Chat better for coding?

They score almost the same on coding (43.6 vs 43.5); test both on your own repository before choosing.

How many benchmarks do Inkling-Small and Longcat Flash Chat share?

17 benchmarks have published results for both models. Inkling-Small has 33 scored results on Noometry and Longcat Flash Chat has 19.

Related comparisons

Go deeper