Model comparison

Claude Fable 5.1 vs Kimi K2.5

Claude Fable 5.1 is the stronger model overall, scoring 69.0 to 48.1 on the Noometry Index. Kimi K2.5 costs 22× less per token, which makes it the better buy when Claude Fable 5.1's lead doesn't matter for your workload.

Last verified . 34 shared benchmarks.

Claude Fable 5.1 Anthropic

69.0

Rank #2 Confirmed

Kimi K2.5 Moonshot AI

48.1

Rank #57 Confirmed

Summary

  • They share 34 benchmarks with published results for both. Claude Fable 5.1 scores higher in 9 categories and Kimi K2.5 in 1 category; 10 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Claude Fable 5.1 leads 76.7 to 31.2.
  • The biggest single-benchmark swing is ARC-AGI-2: 90% for Claude Fable 5.1 and 11.8% for Kimi K2.5.
  • Kimi K2.5 is cheaper at $0.45 / $2.25 per million input/output tokens, against $10 / $50 for Claude Fable 5.1.
  • Claude Fable 5.1 accepts more context: 1M tokens versus 262K.
  • Kimi K2.5 has downloadable open weights; the other is API-only.

Side by side

Claude Fable 5.1 and Kimi K2.5 specifications
Claude Fable 5.1Kimi K2.5
ProviderAnthropicMoonshot AI
Noometry Index69.048.1
Released2026-09-012026-01-27
WeightsProprietaryOpen
Context window1M262K
Max output128K262K
Input $ / M tokens$10$0.45
Output $ / M tokens$50$2.25
Results tracked5251

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Claude Fable 5.1 leads

Claude Fable 5.1: 74.7 (#1), Kimi K2.5: 48.8 (#53)

Coding benchmarks
BenchmarkClaude Fable 5.1Kimi K2.5
LMArena WebDev17441437
SciCode63.1%49%
WeirdML92.9%45.6%
LMArena Coding15281474
ALE-Bench2,143821.65
SWE-bench Verified—73.8%
FrontierCode50.9%—
SWE-bench Verified (bash only)—70.8%
CursorBench51.8%—
SWE-bench Multilingual—67.3%
FrontierSWE56.3%—
GSO88.2%—
MirrorCode73.3%—

Agentic & Tool Use Claude Fable 5.1 leads

Claude Fable 5.1: 50.7 (#5), Kimi K2.5: 34.2 (#48)

Agentic & Tool Use benchmarks
BenchmarkClaude Fable 5.1Kimi K2.5
Vending-Bench 25,4221,198
Terminal-Bench—43.2%
APEX-Agents68.6%—
Remote Labor Index17.9%—
OSWorld—63.3%
GDP.pdf29.6%—

Reasoning Claude Fable 5.1 leads

Claude Fable 5.1: 76.7 (#7), Kimi K2.5: 31.2 (#80)

Reasoning benchmarks
BenchmarkClaude Fable 5.1Kimi K2.5
ARC-AGI-290%11.8%
NYT Connections (extended)90%69.9%
ARC-AGI-197.5%65.3%
CritPt31.1%3.1%
Chess Puzzles47%12%
LMArena Hard Prompts15261453
Epoch Capabilities Index164.7148.03
SimpleBench—46.8%
Kagi LLM Benchmark—78.5%
EnigmaEval—3.4%
Thematic Generalization—69.4%
EBR-Bench57.1%—
Mystery Game Puzzles58%—
DTBench97.6%—
LMCA65.5%—

Math Claude Fable 5.1 leads

Claude Fable 5.1: 89.6 (#4), Kimi K2.5: 51.8 (#53)

Knowledge Claude Fable 5.1 leads

Claude Fable 5.1: 69.6 (#6), Kimi K2.5: 53.6 (#56)

Knowledge benchmarks
BenchmarkClaude Fable 5.1Kimi K2.5
Humanity's Last Exam46.5%24.4%
SimpleQA Verified70.8%34.3%
LMArena Expert15351466
GPQA Diamond—87.6%
Vectara Hallucination Rate—14.2%

Multimodal Claude Fable 5.1 leads

Claude Fable 5.1: 53.9 (#4), Kimi K2.5: 41.1 (#39)

Multimodal benchmarks
BenchmarkClaude Fable 5.1Kimi K2.5
LMArena Vision13181269
LMArena Document15131430
Blueprint-Bench 241.9%—
Furniture Assembly70%—

Multilingual Claude Fable 5.1 leads

Claude Fable 5.1: 59.1 (#3), Kimi K2.5: 53.9 (#53)

Multilingual benchmarks
BenchmarkClaude Fable 5.1Kimi K2.5
LMArena Non-English15071433
LMArena Chinese15861495
LMArena French15251454
LMArena German15001441
LMArena Japanese15431421
LMArena Korean15341410
LMArena Russian15211435
LMArena Spanish15161450

Instruction Following Claude Fable 5.1 leads

Claude Fable 5.1: 79.2 (#6), Kimi K2.5: 75.3 (#64)

Instruction Following benchmarks
BenchmarkClaude Fable 5.1Kimi K2.5
LMArena Instruction Following15171431

Long Context Kimi K2.5 leads

Claude Fable 5.1: 46.7 (#20), Kimi K2.5: 52.1 (#7)

Long Context benchmarks
BenchmarkClaude Fable 5.1Kimi K2.5
LMArena Longer Query15221445
Fiction.LiveBench—86.1%
CL-bench—19.3%
CL-bench Life—13.2%

Writing & Preference Claude Fable 5.1 leads

Claude Fable 5.1: 79.2 (#2), Kimi K2.5: 65.1 (#53)

Writing & Preference benchmarks
BenchmarkClaude Fable 5.1Kimi K2.5
LMArena Text15101445
LMArena Creative Writing15071423
EQ-Bench Creative Writing21621579
LMArena Multi-Turn14921444

Frequently asked questions

Is Claude Fable 5.1 better than Kimi K2.5?

Claude Fable 5.1 is the stronger model overall, scoring 69.0 to 48.1 on the Noometry Index. Kimi K2.5 costs 22× less per token, which makes it the better buy when Claude Fable 5.1's lead doesn't matter for your workload.

Which is cheaper, Claude Fable 5.1 or Kimi K2.5?

Kimi K2.5 is cheaper. It lists at $0.45 per million input tokens and $2.25 per million output tokens; Claude Fable 5.1 lists at $10 and $50.

Is Claude Fable 5.1 or Kimi K2.5 better for coding?

Claude Fable 5.1 scores higher on coding benchmarks: 74.7 versus 48.8 in the Noometry coding category.

Which has the bigger context window?

Claude Fable 5.1 does, with 1M tokens against 262K.

How many benchmarks do Claude Fable 5.1 and Kimi K2.5 share?

34 benchmarks have published results for both models. Claude Fable 5.1 has 52 scored results on Noometry and Kimi K2.5 has 51.

Related comparisons

Go deeper