Moonshot AI, open weights

Kimi K2.5

Kimi K2.5 by Moonshot AI ranks 57th of 354 ranked models on the Noometry Index as of October 2026, with a score of 48.1. Its strongest category is long context, where it ranks 7th. API pricing starts at $0.45 per million input tokens and $2.25 per million output tokens, with a 262K-token context window.

Last verified

Specifications

Noometry rank
#57 of 354
Index score
48.1
Evidence
Confirmed 51 results
Provider
Moonshot AI
Released
January 27, 2026
Weights
Open weights
Reasoning
Yes
Context window
262K
Max output
262K
Input price
$0.45 / M
Output price
$2.25 / M
Blended price
$0.90 / M
Output speed
66 tokens/s Kagi
Value
#93 of 219
Knowledge cutoff
January 2025
Input
text, image

Category scores

Each category score combines every public result we have in that category.

Kimi K2.5 category scores
  1. Coding 48.8
  2. Agentic & Tool Use 34.2
  3. Reasoning 31.2
  4. Math 51.8
  5. Knowledge 53.6
  6. Multimodal 41.1
  7. Multilingual 53.9
  8. Instruction Following 75.3
  9. Long Context 52.1
  10. Writing & Preference 65.1
Kimi K2.5 category ranks
CategoryScoreRankResults
Coding48.8#537
Agentic & Tool Use34.2#482
Reasoning31.2#8010
Math51.8#533
Knowledge53.6#565
Multimodal41.1#391
Multilingual53.9#531
Instruction Following75.3#641
Long Context52.1#74
Writing & Preference65.1#534

Strengths and weaknesses

Categories where Kimi K2.5 places highest and lowest among the models ranked in each, with its score against that category's median.

Strongest categories

Kimi K2.5: strongest categories
CategoryScorevs medianRank
Long Context52.1+11.2#7 of 296, top 3%
Coding48.8+10.1#53 of 340, top 16%
Math51.8+15.3#53 of 327, top 17%

Weakest categories

Kimi K2.5: weakest categories
CategoryScorevs medianRank
Agentic & Tool Use34.2+3.8#48 of 154, top 32%
Multimodal41.1+2.6#39 of 128, top 31%
Reasoning31.2+7.6#80 of 350, top 23%

Closest competitors

The models ranked just above and below Kimi K2.5. When scores are this close, price and speed are often the better way to choose.

Models ranked closest to Kimi K2.5
ModelRankScoreBlended $/MSpeed
GPT-5.1#5349.0$3.44—Compare
Grok 4.20 (Non-Reasoning)#5448.6$1.5661Compare
MiMo-V2.6-Flash#5548.5$0.18—Compare
Grok 4#5648.1—1Compare
Step 5 Preview#5847.9$1.43—Compare
GLM-5.1#5947.8$2.15—Compare
Kimi K2.6#6047.7$1.71—Compare
o3#6147.5$3.503Compare

Sponsored placements are available on pages like this one. Advertise on Noometry

Benchmark results

Every published result we track, with its source. Bold rows are the ones used for ranking; where several exist we prefer independent runs over self-reported numbers.

Coding

Kimi K2.5 Coding benchmark results
BenchmarkScorePositionSettingSourceDate
SWE-bench Verified73.8%#18 of 32, top 57%Epoch AI2026-02-17
SWE-bench Verified (bash only)70.8%#10 of 39, top 26%highSWE-bench2026-02-17
LMArena WebDev1437#61 of 113, top 54%thinkingLMArena2026-10-08
SWE-bench Multilingual67.3%#7 of 13, top 54%SWE-bench2026-02-13
SciCode49%#46 of 121, top 39%Epoch AI
WeirdML45.6%#60 of 119, top 51%Epoch AI
LMArena Coding1474#51 of 294, top 18%thinkingLMArena2026-10-08
ALE-Bench821.65#55 of 105, top 53%Epoch AI

Agentic & Tool Use

Kimi K2.5 Agentic & Tool Use benchmark results
BenchmarkScorePositionSettingSourceDate
Terminal-Bench43.2%#22 of 41, top 54%Epoch AI
OSWorld63.3%#3 of 8, top 38%Epoch AI
Vending-Bench 21,198#45 of 60, top 75%Epoch AI

Reasoning

Kimi K2.5 Reasoning benchmark results
BenchmarkScorePositionSettingSourceDate
ARC-AGI-211.8%#47 of 83, top 57%Epoch AI
SimpleBench46.8%#45 of 77, top 59%Epoch AI
Kagi LLM Benchmark78.5%#9 of 99, top 10%Kagi LLM Benchmark
Kagi LLM Benchmark63.8%Kagi LLM Benchmark
NYT Connections (extended)69.9%#49 of 91, top 54%Lech Mazur benchmarks
ARC-AGI-165.3%#46 of 83, top 56%Epoch AI
CritPt3.1%#65 of 134, top 49%Epoch AI
Chess Puzzles12%#78 of 129, top 61%Epoch AI2026-01-28
EnigmaEval3.4%#25 of 38, top 66%Epoch AI
Thematic Generalization69.4%#7 of 23, top 31%Lech Mazur benchmarks
LMArena Hard Prompts1453#55 of 297, top 19%thinkingLMArena2026-10-08
Epoch Capabilities Index148.03#62 of 213, top 30%Epoch AI2026-01-27

Math

Kimi K2.5 Math benchmark results
BenchmarkScorePositionSettingSourceDate
MathArena Final-Answer Competitions62.3%#21 of 29, top 73%thinkMathArena
OTIS Mock AIME 2024-202592.2%#46 of 173, top 27%Epoch AI2026-02-02
LMArena Math1470#40 of 285, top 15%thinkingLMArena2026-10-08
FrontierMath (Feb 2025 set)27.9%#21 of 68, top 31%Epoch AI2026-02-03
FrontierMath Tier 4 (v1)4.2%#26 of 55, top 48%Epoch AI2026-02-02

Knowledge

Kimi K2.5 Knowledge benchmark results
BenchmarkScorePositionSettingSourceDate
GPQA Diamond87.6%#53 of 186, top 29%Epoch AI2026-02-02
Humanity's Last Exam24.4%#15 of 41, top 37%Epoch AI
SimpleQA Verified34.3%#49 of 77, top 64%Epoch AI2026-08-27
Vectara Hallucination Rate (lower is better)14.2%#84 of 96, top 88%Vectara Hallucination Leaderboard
LMArena Expert1466#53 of 273, top 20%thinkingLMArena2026-10-08

Multimodal

Kimi K2.5 Multimodal benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Vision1269#37 of 122, top 31%thinkingLMArena2026-10-09
LMArena Document1430#29 of 38, top 77%thinkingLMArena2026-09-13

Multilingual

Kimi K2.5 Multilingual benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Non-English1433#53 of 297, top 18%thinkingLMArena2026-10-08
LMArena Chinese1495#52 of 285, top 19%thinkingLMArena2026-10-08
LMArena French1454#65 of 223, top 30%thinkingLMArena2026-10-08
LMArena German1441#54 of 231, top 24%thinkingLMArena2026-10-08
LMArena Japanese1421#39 of 211, top 19%thinkingLMArena2026-10-08
LMArena Korean1410#44 of 213, top 21%thinkingLMArena2026-10-08
LMArena Russian1435#59 of 283, top 21%thinkingLMArena2026-10-08
LMArena Spanish1450#51 of 226, top 23%thinkingLMArena2026-10-08

Instruction Following

Kimi K2.5 Instruction Following benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Instruction Following1431#60 of 298, top 21%thinkingLMArena2026-10-08

Long Context

Kimi K2.5 Long Context benchmark results
BenchmarkScorePositionSettingSourceDate
Fiction.LiveBench86.1%#7 of 47, top 15%Epoch AI
CL-bench19.3%#9 of 19, top 48%Epoch AI
CL-bench Life13.2%#7 of 13, top 54%Epoch AI
LMArena Longer Query1445#58 of 291, top 20%thinkingLMArena2026-10-08

Writing & Preference

Kimi K2.5 Writing & Preference benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Text1445#52 of 297, top 18%thinkingLMArena2026-10-08
LMArena Creative Writing1423#52 of 295, top 18%thinkingLMArena2026-10-08
EQ-Bench Creative Writing1579#46 of 115, top 40%EQ-Bench
LMArena Multi-Turn1444#67 of 295, top 23%thinkingLMArena2026-10-08

API pricing by provider

Kimi K2.5 API prices
RouteInput $/MOutput $/MCached input $/MChecked
azure$0.60$3—2026-10-10
bedrock$0.60$3—2026-10-10
deepinfra$0.45$2.25$0.072026-10-10
openrouter$0.50$2.50$0.152026-10-10

Compare Kimi K2.5

Other Moonshot AI models

Frequently asked questions

How good is Kimi K2.5?

Kimi K2.5 by Moonshot AI ranks 57th of 354 ranked models on the Noometry Index as of October 2026, with a score of 48.1. Its strongest category is long context, where it ranks 7th. API pricing starts at $0.45 per million input tokens and $2.25 per million output tokens, with a 262K-token context window.

How much does Kimi K2.5 cost?

Kimi K2.5 costs $0.45 per million input tokens and $2.25 per million output tokens on deepinfra, with cached input at $0.07.

What is Kimi K2.5's context window?

Kimi K2.5 accepts up to 262K tokens of input and can write up to 262K tokens in one response.

Is Kimi K2.5 open source?

Yes. Kimi K2.5's weights are downloadable from Hugging Face (moonshotai/Kimi-K2.5); check the license for commercial terms.

How fast is Kimi K2.5?

Kimi K2.5 generated about 66 output tokens per second in the Kagi LLM Benchmark's timed runs. Speed varies by provider, load and reasoning effort.

What are Kimi K2.5's strengths and weaknesses?

Relative to other ranked models, Kimi K2.5 places best in long context, coding, math and lowest in agentic & tool use, multimodal, reasoning.

What is Kimi K2.5 best at?

Its best category is long context, where it ranks 7th on Noometry.