Model comparison

Claude Fable 5.1 vs Kimi K2.6

Claude Fable 5.1 is the stronger model overall, scoring 69.0 to 47.7 on the Noometry Index. Kimi K2.6 costs 12× less per token, which makes it the better buy when Claude Fable 5.1's lead doesn't matter for your workload.

Last verified . 41 shared benchmarks.

Claude Fable 5.1 Anthropic

69.0

Rank #2 Confirmed

Kimi K2.6 Moonshot AI

47.7

Rank #60 Confirmed

Summary

  • They share 41 benchmarks with published results for both. Claude Fable 5.1 scores higher in 10 categories and Kimi K2.6 in 0 categories; 10 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Claude Fable 5.1 leads 76.7 to 40.5.
  • The biggest single-benchmark swing is ProofBench: 100% for Claude Fable 5.1 and 16% for Kimi K2.6.
  • Kimi K2.6 is cheaper at $0.95 / $4 per million input/output tokens, against $10 / $50 for Claude Fable 5.1.
  • Claude Fable 5.1 accepts more context: 1M tokens versus 262K.
  • Kimi K2.6 has downloadable open weights; the other is API-only.

Side by side

Claude Fable 5.1 and Kimi K2.6 specifications
Claude Fable 5.1Kimi K2.6
ProviderAnthropicMoonshot AI
Noometry Index69.047.7
Released2026-09-012026-04-20
WeightsProprietaryOpen
Context window1M262K
Max output128K262K
Input $ / M tokens$10$0.95
Output $ / M tokens$50$4
Results tracked5251

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Claude Fable 5.1 leads

Claude Fable 5.1: 74.7 (#1), Kimi K2.6: 50.7 (#43)

Coding benchmarks
BenchmarkClaude Fable 5.1Kimi K2.6
LMArena WebDev17441509
SciCode63.1%53.5%
WeirdML92.9%55.9%
LMArena Coding15281488
ALE-Bench2,1431,093
SWE-bench Verified—76.7%
FrontierCode50.9%—
CursorBench51.8%—
FrontierSWE56.3%—
GSO88.2%—
MirrorCode73.3%—

Agentic & Tool Use Claude Fable 5.1 leads

Claude Fable 5.1: 50.7 (#5), Kimi K2.6: 21.9 (#137)

Agentic & Tool Use benchmarks
BenchmarkClaude Fable 5.1Kimi K2.6
GDP.pdf29.6%12%
Vending-Bench 25,4226,205
APEX-Agents68.6%—
OSWorld 2.0—4.6%
Remote Labor Index17.9%—
ExploitBench—18.4%
GBAEval—0.9%

Reasoning Claude Fable 5.1 leads

Claude Fable 5.1: 76.7 (#7), Kimi K2.6: 40.5 (#55)

Reasoning benchmarks
BenchmarkClaude Fable 5.1Kimi K2.6
NYT Connections (extended)90%87.2%
CritPt31.1%8%
Chess Puzzles47%26%
EBR-Bench57.1%2.4%
LMArena Hard Prompts15261470
Mystery Game Puzzles58%18%
DTBench97.6%90.9%
LMCA65.5%37.3%
Epoch Capabilities Index164.7151.05
ARC-AGI-290%—
ARC-AGI-197.5%—

Math Claude Fable 5.1 leads

Claude Fable 5.1: 89.6 (#4), Kimi K2.6: 57.0 (#41)

Knowledge Claude Fable 5.1 leads

Claude Fable 5.1: 69.6 (#6), Kimi K2.6: 54.0 (#54)

Knowledge benchmarks
BenchmarkClaude Fable 5.1Kimi K2.6
SimpleQA Verified70.8%34.9%
LMArena Expert15351491
GPQA Diamond—90.8%
Humanity's Last Exam46.5%—
Vectara Hallucination Rate—10.8%

Multimodal Claude Fable 5.1 leads

Claude Fable 5.1: 53.9 (#4), Kimi K2.6: 31.6 (#103)

Multimodal benchmarks
BenchmarkClaude Fable 5.1Kimi K2.6
LMArena Vision13181283
Blueprint-Bench 241.9%3.9%
Furniture Assembly70%21.7%
LMArena Document15131451

Multilingual Claude Fable 5.1 leads

Claude Fable 5.1: 59.1 (#3), Kimi K2.6: 54.9 (#37)

Multilingual benchmarks
BenchmarkClaude Fable 5.1Kimi K2.6
LMArena Non-English15071446
LMArena Chinese15861521
LMArena French15251471
LMArena German15001450
LMArena Japanese15431443
LMArena Korean15341427
LMArena Russian15211446
LMArena Spanish15161464

Instruction Following Claude Fable 5.1 leads

Claude Fable 5.1: 79.2 (#6), Kimi K2.6: 76.3 (#43)

Instruction Following benchmarks
BenchmarkClaude Fable 5.1Kimi K2.6
LMArena Instruction Following15171451

Long Context Claude Fable 5.1 leads

Claude Fable 5.1: 46.7 (#20), Kimi K2.6: 44.9 (#52)

Long Context benchmarks
BenchmarkClaude Fable 5.1Kimi K2.6
LMArena Longer Query15221468

Writing & Preference Claude Fable 5.1 leads

Claude Fable 5.1: 79.2 (#2), Kimi K2.6: 68.5 (#26)

Writing & Preference benchmarks
BenchmarkClaude Fable 5.1Kimi K2.6
LMArena Text15101455
LMArena Creative Writing15071434
EQ-Bench Creative Writing21621725
LMArena Multi-Turn14921453
EQ-Bench 4—1202

Frequently asked questions

Is Claude Fable 5.1 better than Kimi K2.6?

Claude Fable 5.1 is the stronger model overall, scoring 69.0 to 47.7 on the Noometry Index. Kimi K2.6 costs 12× less per token, which makes it the better buy when Claude Fable 5.1's lead doesn't matter for your workload.

Which is cheaper, Claude Fable 5.1 or Kimi K2.6?

Kimi K2.6 is cheaper. It lists at $0.95 per million input tokens and $4 per million output tokens; Claude Fable 5.1 lists at $10 and $50.

Is Claude Fable 5.1 or Kimi K2.6 better for coding?

Claude Fable 5.1 scores higher on coding benchmarks: 74.7 versus 50.7 in the Noometry coding category.

Which has the bigger context window?

Claude Fable 5.1 does, with 1M tokens against 262K.

How many benchmarks do Claude Fable 5.1 and Kimi K2.6 share?

41 benchmarks have published results for both models. Claude Fable 5.1 has 52 scored results on Noometry and Kimi K2.6 has 51.

Related comparisons

Go deeper