Moonshot AI, open weights
Kimi K2.5
Kimi K2.5 by Moonshot AI ranks 57th of 354 ranked models on the Noometry Index as of October 2026, with a score of 48.1. Its strongest category is long context, where it ranks 7th. API pricing starts at $0.45 per million input tokens and $2.25 per million output tokens, with a 262K-token context window.
Last verified
Specifications
- Noometry rank
- #57 of 354
- Index score
- 48.1
- Evidence
- Confirmed 51 results
- Provider
- Moonshot AI
- Released
- January 27, 2026
- Weights
- Open weights
- Reasoning
- Yes
- Context window
- 262K
- Max output
- 262K
- Input price
- $0.45 / M
- Output price
- $2.25 / M
- Blended price
- $0.90 / M
- Output speed
- 66 tokens/s Kagi
- Value
- #93 of 219
- Knowledge cutoff
- January 2025
- Input
- text, image
- Hugging Face
- moonshotai/Kimi-K2.5
Category scores
Each category score combines every public result we have in that category.
- Coding 48.8
- Agentic & Tool Use 34.2
- Reasoning 31.2
- Math 51.8
- Knowledge 53.6
- Multimodal 41.1
- Multilingual 53.9
- Instruction Following 75.3
- Long Context 52.1
- Writing & Preference 65.1
| Category | Score | Rank | Results |
|---|---|---|---|
| Coding | 48.8 | #53 | 7 |
| Agentic & Tool Use | 34.2 | #48 | 2 |
| Reasoning | 31.2 | #80 | 10 |
| Math | 51.8 | #53 | 3 |
| Knowledge | 53.6 | #56 | 5 |
| Multimodal | 41.1 | #39 | 1 |
| Multilingual | 53.9 | #53 | 1 |
| Instruction Following | 75.3 | #64 | 1 |
| Long Context | 52.1 | #7 | 4 |
| Writing & Preference | 65.1 | #53 | 4 |
Strengths and weaknesses
Categories where Kimi K2.5 places highest and lowest among the models ranked in each, with its score against that category's median.
Strongest categories
| Category | Score | vs median | Rank |
|---|---|---|---|
| Long Context | 52.1 | +11.2 | #7 of 296, top 3% |
| Coding | 48.8 | +10.1 | #53 of 340, top 16% |
| Math | 51.8 | +15.3 | #53 of 327, top 17% |
Weakest categories
| Category | Score | vs median | Rank |
|---|---|---|---|
| Agentic & Tool Use | 34.2 | +3.8 | #48 of 154, top 32% |
| Multimodal | 41.1 | +2.6 | #39 of 128, top 31% |
| Reasoning | 31.2 | +7.6 | #80 of 350, top 23% |
Closest competitors
The models ranked just above and below Kimi K2.5. When scores are this close, price and speed are often the better way to choose.
| Model | Rank | Score | Blended $/M | Speed | |
|---|---|---|---|---|---|
| GPT-5.1 | #53 | 49.0 | $3.44 | — | Compare |
| Grok 4.20 (Non-Reasoning) | #54 | 48.6 | $1.56 | 61 | Compare |
| MiMo-V2.6-Flash | #55 | 48.5 | $0.18 | — | Compare |
| Grok 4 | #56 | 48.1 | — | 1 | Compare |
| Step 5 Preview | #58 | 47.9 | $1.43 | — | Compare |
| GLM-5.1 | #59 | 47.8 | $2.15 | — | Compare |
| Kimi K2.6 | #60 | 47.7 | $1.71 | — | Compare |
| o3 | #61 | 47.5 | $3.50 | 3 | Compare |
Sponsored placements are available on pages like this one. Advertise on Noometry
Benchmark results
Every published result we track, with its source. Bold rows are the ones used for ranking; where several exist we prefer independent runs over self-reported numbers.
Coding
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| SWE-bench Verified | 73.8% | #18 of 32, top 57% | Epoch AI | 2026-02-17 | |
| SWE-bench Verified (bash only) | 70.8% | #10 of 39, top 26% | high | SWE-bench | 2026-02-17 |
| LMArena WebDev | 1437 | #61 of 113, top 54% | thinking | LMArena | 2026-10-08 |
| SWE-bench Multilingual | 67.3% | #7 of 13, top 54% | SWE-bench | 2026-02-13 | |
| SciCode | 49% | #46 of 121, top 39% | Epoch AI | ||
| WeirdML | 45.6% | #60 of 119, top 51% | Epoch AI | ||
| LMArena Coding | 1474 | #51 of 294, top 18% | thinking | LMArena | 2026-10-08 |
| ALE-Bench | 821.65 | #55 of 105, top 53% | Epoch AI |
Agentic & Tool Use
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| Terminal-Bench | 43.2% | #22 of 41, top 54% | Epoch AI | ||
| OSWorld | 63.3% | #3 of 8, top 38% | Epoch AI | ||
| Vending-Bench 2 | 1,198 | #45 of 60, top 75% | Epoch AI |
Reasoning
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| ARC-AGI-2 | 11.8% | #47 of 83, top 57% | Epoch AI | ||
| SimpleBench | 46.8% | #45 of 77, top 59% | Epoch AI | ||
| Kagi LLM Benchmark | 78.5% | #9 of 99, top 10% | Kagi LLM Benchmark | ||
| Kagi LLM Benchmark | 63.8% | Kagi LLM Benchmark | |||
| NYT Connections (extended) | 69.9% | #49 of 91, top 54% | Lech Mazur benchmarks | ||
| ARC-AGI-1 | 65.3% | #46 of 83, top 56% | Epoch AI | ||
| CritPt | 3.1% | #65 of 134, top 49% | Epoch AI | ||
| Chess Puzzles | 12% | #78 of 129, top 61% | Epoch AI | 2026-01-28 | |
| EnigmaEval | 3.4% | #25 of 38, top 66% | Epoch AI | ||
| Thematic Generalization | 69.4% | #7 of 23, top 31% | Lech Mazur benchmarks | ||
| LMArena Hard Prompts | 1453 | #55 of 297, top 19% | thinking | LMArena | 2026-10-08 |
| Epoch Capabilities Index | 148.03 | #62 of 213, top 30% | Epoch AI | 2026-01-27 |
Math
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| MathArena Final-Answer Competitions | 62.3% | #21 of 29, top 73% | think | MathArena | |
| OTIS Mock AIME 2024-2025 | 92.2% | #46 of 173, top 27% | Epoch AI | 2026-02-02 | |
| LMArena Math | 1470 | #40 of 285, top 15% | thinking | LMArena | 2026-10-08 |
| FrontierMath (Feb 2025 set) | 27.9% | #21 of 68, top 31% | Epoch AI | 2026-02-03 | |
| FrontierMath Tier 4 (v1) | 4.2% | #26 of 55, top 48% | Epoch AI | 2026-02-02 |
Knowledge
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| GPQA Diamond | 87.6% | #53 of 186, top 29% | Epoch AI | 2026-02-02 | |
| Humanity's Last Exam | 24.4% | #15 of 41, top 37% | Epoch AI | ||
| SimpleQA Verified | 34.3% | #49 of 77, top 64% | Epoch AI | 2026-08-27 | |
| Vectara Hallucination Rate (lower is better) | 14.2% | #84 of 96, top 88% | Vectara Hallucination Leaderboard | ||
| LMArena Expert | 1466 | #53 of 273, top 20% | thinking | LMArena | 2026-10-08 |
Multimodal
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Vision | 1269 | #37 of 122, top 31% | thinking | LMArena | 2026-10-09 |
| LMArena Document | 1430 | #29 of 38, top 77% | thinking | LMArena | 2026-09-13 |
Multilingual
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Non-English | 1433 | #53 of 297, top 18% | thinking | LMArena | 2026-10-08 |
| LMArena Chinese | 1495 | #52 of 285, top 19% | thinking | LMArena | 2026-10-08 |
| LMArena French | 1454 | #65 of 223, top 30% | thinking | LMArena | 2026-10-08 |
| LMArena German | 1441 | #54 of 231, top 24% | thinking | LMArena | 2026-10-08 |
| LMArena Japanese | 1421 | #39 of 211, top 19% | thinking | LMArena | 2026-10-08 |
| LMArena Korean | 1410 | #44 of 213, top 21% | thinking | LMArena | 2026-10-08 |
| LMArena Russian | 1435 | #59 of 283, top 21% | thinking | LMArena | 2026-10-08 |
| LMArena Spanish | 1450 | #51 of 226, top 23% | thinking | LMArena | 2026-10-08 |
Instruction Following
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Instruction Following | 1431 | #60 of 298, top 21% | thinking | LMArena | 2026-10-08 |
Long Context
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| Fiction.LiveBench | 86.1% | #7 of 47, top 15% | Epoch AI | ||
| CL-bench | 19.3% | #9 of 19, top 48% | Epoch AI | ||
| CL-bench Life | 13.2% | #7 of 13, top 54% | Epoch AI | ||
| LMArena Longer Query | 1445 | #58 of 291, top 20% | thinking | LMArena | 2026-10-08 |
Writing & Preference
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Text | 1445 | #52 of 297, top 18% | thinking | LMArena | 2026-10-08 |
| LMArena Creative Writing | 1423 | #52 of 295, top 18% | thinking | LMArena | 2026-10-08 |
| EQ-Bench Creative Writing | 1579 | #46 of 115, top 40% | EQ-Bench | ||
| LMArena Multi-Turn | 1444 | #67 of 295, top 23% | thinking | LMArena | 2026-10-08 |
API pricing by provider
| Route | Input $/M | Output $/M | Cached input $/M | Checked |
|---|---|---|---|---|
| azure | $0.60 | $3 | — | 2026-10-10 |
| bedrock | $0.60 | $3 | — | 2026-10-10 |
| deepinfra | $0.45 | $2.25 | $0.07 | 2026-10-10 |
| openrouter | $0.50 | $2.50 | $0.15 | 2026-10-10 |
Compare Kimi K2.5
- Kimi K2.5 vs Grok 4
- Kimi K2.5 vs Step 5 Preview
- Kimi K2.5 vs MiMo-V2.6-Flash
- Kimi K2.5 vs GLM-5.1
- Kimi K2.5 vs Grok 4.20 (Non-Reasoning)
- Kimi K2.5 vs Kimi K2.6
- Kimi K2.5 vs GPT-6 Astra
- Kimi K2.5 vs Claude Fable 5.1
- Kimi K2.5 vs Gemini 3.8 Flash
- Kimi K2.5 vs Grok 4.6
- Kimi K2.5 vs Qwen3.8 Max
- Kimi K2.5 vs GLM-5.3
- Kimi K2.5 vs Muse Spark 1.3
- Kimi K2.5 vs DeepSeek V4 Pro
Other Moonshot AI models
Frequently asked questions
How good is Kimi K2.5?
Kimi K2.5 by Moonshot AI ranks 57th of 354 ranked models on the Noometry Index as of October 2026, with a score of 48.1. Its strongest category is long context, where it ranks 7th. API pricing starts at $0.45 per million input tokens and $2.25 per million output tokens, with a 262K-token context window.
How much does Kimi K2.5 cost?
Kimi K2.5 costs $0.45 per million input tokens and $2.25 per million output tokens on deepinfra, with cached input at $0.07.
What is Kimi K2.5's context window?
Kimi K2.5 accepts up to 262K tokens of input and can write up to 262K tokens in one response.
Is Kimi K2.5 open source?
Yes. Kimi K2.5's weights are downloadable from Hugging Face (moonshotai/Kimi-K2.5); check the license for commercial terms.
How fast is Kimi K2.5?
Kimi K2.5 generated about 66 output tokens per second in the Kagi LLM Benchmark's timed runs. Speed varies by provider, load and reasoning effort.
What are Kimi K2.5's strengths and weaknesses?
Relative to other ranked models, Kimi K2.5 places best in long context, coding, math and lowest in agentic & tool use, multimodal, reasoning.
What is Kimi K2.5 best at?
Its best category is long context, where it ranks 7th on Noometry.