Moonshot AI, open weights

# Kimi K2.5

> Kimi K2.5 by Moonshot AI, released January 2026. Ranked #57 of 354 with a Noometry Index of 48.1. API: $0.45 in / $2.25 out per M tokens. 262K context. Scores, sources and comparisons.
- Canonical page: https://noometry.com/models/kimi-k2-5
- Last updated: 2026-10-10
- Title: Kimi K2.5 Benchmarks, Price & Rank (October 2026) | Noometry

Kimi K2.5 by Moonshot AI ranks 57th of 354 ranked models on the Noometry Index as of October 2026, with a score of 48.1. Its strongest category is long context, where it ranks 7th. API pricing starts at $0.45 per million input tokens and $2.25 per million output tokens, with a 262K-token context window.

Last verified October 10, 2026

## Specifications

- **Noometry rank:** #57 of 354
- **Index score:** 48.1
- **Evidence:** Confirmed 51 results
- **Provider:** [Moonshot AI](https://noometry.com/providers/moonshot)
- **Released:** January 27, 2026
- **Weights:** Open weights
- **Reasoning:** Yes
- **Context window:** 262K
- **Max output:** 262K
- **Input price:** $0.45 / M
- **Output price:** $2.25 / M
- **Blended price:** $0.90 / M
- **Output speed:** 66 tokens/s [Kagi](https://help.kagi.com/kagi/ai/llm-benchmark.html)
- **Value:** #93 of 219
- **Knowledge cutoff:** January 2025
- **Input:** text, image
- **Hugging Face:** [moonshotai/Kimi-K2.5](https://huggingface.co/moonshotai/Kimi-K2.5)

## Category scores

Each category score combines every public result we have in that category.

Kimi K2.5 category scores

1.  Coding 48.8
2.  Agentic & Tool Use 34.2
3.  Reasoning 31.2
4.  Math 51.8
5.  Knowledge 53.6
6.  Multimodal 41.1
7.  Multilingual 53.9
8.  Instruction Following 75.3
9.  Long Context 52.1
10.  Writing & Preference 65.1
11.  020406080

Kimi K2.5 category ranks
| Category | Score | Rank | Results |
| --- | --- | --- | --- |
| [Coding](https://noometry.com/best/coding) | 48.8 | #53 | 7 |
| [Agentic & Tool Use](https://noometry.com/best/agentic) | 34.2 | #48 | 2 |
| [Reasoning](https://noometry.com/best/reasoning) | 31.2 | #80 | 10 |
| [Math](https://noometry.com/best/math) | 51.8 | #53 | 3 |
| [Knowledge](https://noometry.com/best/knowledge) | 53.6 | #56 | 5 |
| [Multimodal](https://noometry.com/best/multimodal) | 41.1 | #39 | 1 |
| [Multilingual](https://noometry.com/best/multilingual) | 53.9 | #53 | 1 |
| [Instruction Following](https://noometry.com/best/instruction-following) | 75.3 | #64 | 1 |
| [Long Context](https://noometry.com/best/long-context) | 52.1 | #7 | 4 |
| [Writing & Preference](https://noometry.com/best/writing) | 65.1 | #53 | 4 |

## Strengths and weaknesses

Categories where Kimi K2.5 places highest and lowest among the models ranked in each, with its score against that category's median.

### Strongest categories

Kimi K2.5: strongest categories
| Category | Score | vs median | Rank |
| --- | --- | --- | --- |
| [Long Context](https://noometry.com/best/long-context) | 52.1 | +11.2 | #7 of 296, top 3% |
| [Coding](https://noometry.com/best/coding) | 48.8 | +10.1 | #53 of 340, top 16% |
| [Math](https://noometry.com/best/math) | 51.8 | +15.3 | #53 of 327, top 17% |

### Weakest categories

Kimi K2.5: weakest categories
| Category | Score | vs median | Rank |
| --- | --- | --- | --- |
| [Agentic & Tool Use](https://noometry.com/best/agentic) | 34.2 | +3.8 | #48 of 154, top 32% |
| [Multimodal](https://noometry.com/best/multimodal) | 41.1 | +2.6 | #39 of 128, top 31% |
| [Reasoning](https://noometry.com/best/reasoning) | 31.2 | +7.6 | #80 of 350, top 23% |

## Closest competitors

The models ranked just above and below Kimi K2.5. When scores are this close, price and speed are often the better way to choose.

Models ranked closest to Kimi K2.5
| Model | Rank | Score | Blended $/M | Speed |  |
| --- | --- | --- | --- | --- | --- |
| [GPT-5.1](https://noometry.com/models/gpt-5-1) | #53 | 49.0 | $3.44 | — | [Compare](https://noometry.com/compare/gpt-5-1-vs-kimi-k2-5) |
| [Grok 4.20 (Non-Reasoning)](https://noometry.com/models/grok-4-20) | #54 | 48.6 | $1.56 | 61 | [Compare](https://noometry.com/compare/grok-4-20-vs-kimi-k2-5) |
| [MiMo-V2.6-Flash](https://noometry.com/models/mimo-v2-6-flash) | #55 | 48.5 | $0.18 | — | [Compare](https://noometry.com/compare/kimi-k2-5-vs-mimo-v2-6-flash) |
| [Grok 4](https://noometry.com/models/grok-4) | #56 | 48.1 | — | 1 | [Compare](https://noometry.com/compare/grok-4-vs-kimi-k2-5) |
| [Step 5 Preview](https://noometry.com/models/step-5-preview) | #58 | 47.9 | $1.43 | — | [Compare](https://noometry.com/compare/kimi-k2-5-vs-step-5-preview) |
| [GLM-5.1](https://noometry.com/models/glm-5-1) | #59 | 47.8 | $2.15 | — | [Compare](https://noometry.com/compare/glm-5-1-vs-kimi-k2-5) |
| [Kimi K2.6](https://noometry.com/models/kimi-k2-6) | #60 | 47.7 | $1.71 | — | [Compare](https://noometry.com/compare/kimi-k2-5-vs-kimi-k2-6) |
| [o3](https://noometry.com/models/o3) | #61 | 47.5 | $3.50 | 3 | [Compare](https://noometry.com/compare/kimi-k2-5-vs-o3) |

Sponsored placements are available on pages like this one. [Advertise on Noometry](https://noometry.com/advertise)

## Benchmark results

Every published result we track, with its source. Bold rows are the ones used for ranking; where several exist we prefer independent runs over self-reported numbers.

### Coding

Kimi K2.5 Coding benchmark results
| Benchmark | Score | Position | Setting | Source | Date |
| --- | --- | --- | --- | --- | --- |
| [SWE-bench Verified](https://noometry.com/benchmarks/swe-bench-verified) | 73.8% | #18 of 32, top 57% |  | [Epoch AI](https://epoch.ai/benchmarks) | 2026-02-17 |
| [SWE-bench Verified (bash only)](https://noometry.com/benchmarks/swe-bench-bash-only) | 70.8% | #10 of 39, top 26% | high | [SWE-bench](https://www.swebench.com/) | 2026-02-17 |
| [LMArena WebDev](https://noometry.com/benchmarks/arena-webdev) | 1437 | #61 of 113, top 54% | thinking | [LMArena](https://lmarena.ai/leaderboard/webdev) | 2026-10-08 |
| [SWE-bench Multilingual](https://noometry.com/benchmarks/swe-bench-multilingual) | 67.3% | #7 of 13, top 54% |  | [SWE-bench](https://www.swebench.com/) | 2026-02-13 |
| [SciCode](https://noometry.com/benchmarks/scicode) | 49% | #46 of 121, top 39% |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| [WeirdML](https://noometry.com/benchmarks/weirdml) | 45.6% | #60 of 119, top 51% |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| [LMArena Coding](https://noometry.com/benchmarks/arena-coding) | 1474 | #51 of 294, top 18% | thinking | [LMArena](https://lmarena.ai/leaderboard/text) | 2026-10-08 |
| [ALE-Bench](https://noometry.com/benchmarks/ale-bench) | 821.65 | #55 of 105, top 53% |  | [Epoch AI](https://epoch.ai/benchmarks) |  |

### Agentic & Tool Use

Kimi K2.5 Agentic & Tool Use benchmark results
| Benchmark | Score | Position | Setting | Source | Date |
| --- | --- | --- | --- | --- | --- |
| [Terminal-Bench](https://noometry.com/benchmarks/terminal-bench) | 43.2% | #22 of 41, top 54% |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| [OSWorld](https://noometry.com/benchmarks/osworld) | 63.3% | #3 of 8, top 38% |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| [Vending-Bench 2](https://noometry.com/benchmarks/vending-bench-2) | 1,198 | #45 of 60, top 75% |  | [Epoch AI](https://epoch.ai/benchmarks) |  |

### Reasoning

Kimi K2.5 Reasoning benchmark results
| Benchmark | Score | Position | Setting | Source | Date |
| --- | --- | --- | --- | --- | --- |
| [ARC-AGI-2](https://noometry.com/benchmarks/arc-agi-2) | 11.8% | #47 of 83, top 57% |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| [SimpleBench](https://noometry.com/benchmarks/simplebench) | 46.8% | #45 of 77, top 59% |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| [Kagi LLM Benchmark](https://noometry.com/benchmarks/kagi-reasoning) | 78.5% | #9 of 99, top 10% |  | [Kagi LLM Benchmark](https://help.kagi.com/kagi/ai/llm-benchmark.html) |  |
| [Kagi LLM Benchmark](https://noometry.com/benchmarks/kagi-reasoning) | 63.8% |  |  | [Kagi LLM Benchmark](https://help.kagi.com/kagi/ai/llm-benchmark.html) |  |
| [NYT Connections (extended)](https://noometry.com/benchmarks/nyt-connections) | 69.9% | #49 of 91, top 54% |  | [Lech Mazur benchmarks](https://github.com/lechmazur/nyt-connections) |  |
| [ARC-AGI-1](https://noometry.com/benchmarks/arc-agi-1) | 65.3% | #46 of 83, top 56% |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| [CritPt](https://noometry.com/benchmarks/critpt) | 3.1% | #65 of 134, top 49% |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| [Chess Puzzles](https://noometry.com/benchmarks/chess-puzzles) | 12% | #78 of 129, top 61% |  | [Epoch AI](https://epoch.ai/benchmarks) | 2026-01-28 |
| [EnigmaEval](https://noometry.com/benchmarks/enigmaeval) | 3.4% | #25 of 38, top 66% |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| [Thematic Generalization](https://noometry.com/benchmarks/thematic-generalization) | 69.4% | #7 of 23, top 31% |  | [Lech Mazur benchmarks](https://github.com/lechmazur/generalization) |  |
| [LMArena Hard Prompts](https://noometry.com/benchmarks/arena-hard-prompts) | 1453 | #55 of 297, top 19% | thinking | [LMArena](https://lmarena.ai/leaderboard/text) | 2026-10-08 |
| [Epoch Capabilities Index](https://noometry.com/benchmarks/epoch-capabilities-index) | 148.03 | #62 of 213, top 30% |  | [Epoch AI](https://epoch.ai/eci) | 2026-01-27 |

### Math

Kimi K2.5 Math benchmark results
| Benchmark | Score | Position | Setting | Source | Date |
| --- | --- | --- | --- | --- | --- |
| [MathArena Final-Answer Competitions](https://noometry.com/benchmarks/matharena) | 62.3% | #21 of 29, top 73% | think | [MathArena](https://matharena.ai/) |  |
| [OTIS Mock AIME 2024-2025](https://noometry.com/benchmarks/otis-mock-aime) | 92.2% | #46 of 173, top 27% |  | [Epoch AI](https://epoch.ai/benchmarks) | 2026-02-02 |
| [LMArena Math](https://noometry.com/benchmarks/arena-math) | 1470 | #40 of 285, top 15% | thinking | [LMArena](https://lmarena.ai/leaderboard/text) | 2026-10-08 |
| [FrontierMath (Feb 2025 set)](https://noometry.com/benchmarks/frontiermath-2025-02) | 27.9% | #21 of 68, top 31% |  | [Epoch AI](https://epoch.ai/benchmarks) | 2026-02-03 |
| [FrontierMath Tier 4 (v1)](https://noometry.com/benchmarks/frontiermath-tier-4-v1) | 4.2% | #26 of 55, top 48% |  | [Epoch AI](https://epoch.ai/benchmarks) | 2026-02-02 |

### Knowledge

Kimi K2.5 Knowledge benchmark results
| Benchmark | Score | Position | Setting | Source | Date |
| --- | --- | --- | --- | --- | --- |
| [GPQA Diamond](https://noometry.com/benchmarks/gpqa-diamond) | 87.6% | #53 of 186, top 29% |  | [Epoch AI](https://epoch.ai/benchmarks) | 2026-02-02 |
| [Humanity's Last Exam](https://noometry.com/benchmarks/hle) | 24.4% | #15 of 41, top 37% |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| [SimpleQA Verified](https://noometry.com/benchmarks/simpleqa-verified) | 34.3% | #49 of 77, top 64% |  | [Epoch AI](https://epoch.ai/benchmarks) | 2026-08-27 |
| [Vectara Hallucination Rate](https://noometry.com/benchmarks/vectara-hallucination) (lower is better) | 14.2% | #84 of 96, top 88% |  | [Vectara Hallucination Leaderboard](https://github.com/vectara/hallucination-leaderboard) |  |
| [LMArena Expert](https://noometry.com/benchmarks/arena-expert) | 1466 | #53 of 273, top 20% | thinking | [LMArena](https://lmarena.ai/leaderboard/text) | 2026-10-08 |

### Multimodal

Kimi K2.5 Multimodal benchmark results
| Benchmark | Score | Position | Setting | Source | Date |
| --- | --- | --- | --- | --- | --- |
| [LMArena Vision](https://noometry.com/benchmarks/arena-vision) | 1269 | #37 of 122, top 31% | thinking | [LMArena](https://lmarena.ai/leaderboard/vision) | 2026-10-09 |
| [LMArena Document](https://noometry.com/benchmarks/arena-document) | 1430 | #29 of 38, top 77% | thinking | [LMArena](https://lmarena.ai/leaderboard/document) | 2026-09-13 |

### Multilingual

Kimi K2.5 Multilingual benchmark results
| Benchmark | Score | Position | Setting | Source | Date |
| --- | --- | --- | --- | --- | --- |
| [LMArena Non-English](https://noometry.com/benchmarks/arena-non-english) | 1433 | #53 of 297, top 18% | thinking | [LMArena](https://lmarena.ai/leaderboard/text) | 2026-10-08 |
| [LMArena Chinese](https://noometry.com/benchmarks/arena-chinese) | 1495 | #52 of 285, top 19% | thinking | [LMArena](https://lmarena.ai/leaderboard/text) | 2026-10-08 |
| [LMArena French](https://noometry.com/benchmarks/arena-french) | 1454 | #65 of 223, top 30% | thinking | [LMArena](https://lmarena.ai/leaderboard/text) | 2026-10-08 |
| [LMArena German](https://noometry.com/benchmarks/arena-german) | 1441 | #54 of 231, top 24% | thinking | [LMArena](https://lmarena.ai/leaderboard/text) | 2026-10-08 |
| [LMArena Japanese](https://noometry.com/benchmarks/arena-japanese) | 1421 | #39 of 211, top 19% | thinking | [LMArena](https://lmarena.ai/leaderboard/text) | 2026-10-08 |
| [LMArena Korean](https://noometry.com/benchmarks/arena-korean) | 1410 | #44 of 213, top 21% | thinking | [LMArena](https://lmarena.ai/leaderboard/text) | 2026-10-08 |
| [LMArena Russian](https://noometry.com/benchmarks/arena-russian) | 1435 | #59 of 283, top 21% | thinking | [LMArena](https://lmarena.ai/leaderboard/text) | 2026-10-08 |
| [LMArena Spanish](https://noometry.com/benchmarks/arena-spanish) | 1450 | #51 of 226, top 23% | thinking | [LMArena](https://lmarena.ai/leaderboard/text) | 2026-10-08 |

### Instruction Following

Kimi K2.5 Instruction Following benchmark results
| Benchmark | Score | Position | Setting | Source | Date |
| --- | --- | --- | --- | --- | --- |
| [LMArena Instruction Following](https://noometry.com/benchmarks/arena-instruction-following) | 1431 | #60 of 298, top 21% | thinking | [LMArena](https://lmarena.ai/leaderboard/text) | 2026-10-08 |

### Long Context

Kimi K2.5 Long Context benchmark results
| Benchmark | Score | Position | Setting | Source | Date |
| --- | --- | --- | --- | --- | --- |
| [Fiction.LiveBench](https://noometry.com/benchmarks/fiction-livebench) | 86.1% | #7 of 47, top 15% |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| [CL-bench](https://noometry.com/benchmarks/cl-bench) | 19.3% | #9 of 19, top 48% |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| [CL-bench Life](https://noometry.com/benchmarks/cl-bench-life) | 13.2% | #7 of 13, top 54% |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| [LMArena Longer Query](https://noometry.com/benchmarks/arena-longer-query) | 1445 | #58 of 291, top 20% | thinking | [LMArena](https://lmarena.ai/leaderboard/text) | 2026-10-08 |

### Writing & Preference

Kimi K2.5 Writing & Preference benchmark results
| Benchmark | Score | Position | Setting | Source | Date |
| --- | --- | --- | --- | --- | --- |
| [LMArena Text](https://noometry.com/benchmarks/arena-text) | 1445 | #52 of 297, top 18% | thinking | [LMArena](https://lmarena.ai/leaderboard/text) | 2026-10-08 |
| [LMArena Creative Writing](https://noometry.com/benchmarks/arena-creative-writing) | 1423 | #52 of 295, top 18% | thinking | [LMArena](https://lmarena.ai/leaderboard/text) | 2026-10-08 |
| [EQ-Bench Creative Writing](https://noometry.com/benchmarks/eqbench-creative-writing) | 1579 | #46 of 115, top 40% |  | [EQ-Bench](https://eqbench.com/creative_writing.html) |  |
| [LMArena Multi-Turn](https://noometry.com/benchmarks/arena-multi-turn) | 1444 | #67 of 295, top 23% | thinking | [LMArena](https://lmarena.ai/leaderboard/text) | 2026-10-08 |

## API pricing by provider

Kimi K2.5 API prices
| Route | Input $/M | Output $/M | Cached input $/M | Checked |
| --- | --- | --- | --- | --- |
| [azure](https://learn.microsoft.com/en-us/azure/ai-services/openai/concepts/models) | $0.60 | $3 | — | 2026-10-10 |
| [bedrock](https://docs.aws.amazon.com/bedrock/latest/userguide/models-supported.html) | $0.60 | $3 | — | 2026-10-10 |
| [deepinfra](https://deepinfra.com/models) | $0.45 | $2.25 | $0.07 | 2026-10-10 |
| [openrouter](https://openrouter.ai/moonshotai/kimi-k2.5) | $0.50 | $2.50 | $0.15 | 2026-10-10 |

[All Moonshot AI API prices →](https://noometry.com/llm-pricing/moonshot) [Estimate your cost →](https://noometry.com/tools/cost-calculator)

## Compare Kimi K2.5

-   [Kimi K2.5 vs Grok 4](https://noometry.com/compare/grok-4-vs-kimi-k2-5)
-   [Kimi K2.5 vs Step 5 Preview](https://noometry.com/compare/kimi-k2-5-vs-step-5-preview)
-   [Kimi K2.5 vs MiMo-V2.6-Flash](https://noometry.com/compare/kimi-k2-5-vs-mimo-v2-6-flash)
-   [Kimi K2.5 vs GLM-5.1](https://noometry.com/compare/glm-5-1-vs-kimi-k2-5)
-   [Kimi K2.5 vs Grok 4.20 (Non-Reasoning)](https://noometry.com/compare/grok-4-20-vs-kimi-k2-5)
-   [Kimi K2.5 vs Kimi K2.6](https://noometry.com/compare/kimi-k2-5-vs-kimi-k2-6)
-   [Kimi K2.5 vs GPT-6 Astra](https://noometry.com/compare/gpt-6-astra-vs-kimi-k2-5)
-   [Kimi K2.5 vs Claude Fable 5.1](https://noometry.com/compare/claude-fable-5-1-vs-kimi-k2-5)
-   [Kimi K2.5 vs Gemini 3.8 Flash](https://noometry.com/compare/gemini-3-8-flash-vs-kimi-k2-5)
-   [Kimi K2.5 vs Grok 4.6](https://noometry.com/compare/grok-4-6-vs-kimi-k2-5)
-   [Kimi K2.5 vs Qwen3.8 Max](https://noometry.com/compare/kimi-k2-5-vs-qwen3-8-max)
-   [Kimi K2.5 vs GLM-5.3](https://noometry.com/compare/glm-5-3-vs-kimi-k2-5)
-   [Kimi K2.5 vs Muse Spark 1.3](https://noometry.com/compare/kimi-k2-5-vs-muse-spark-1-3)
-   [Kimi K2.5 vs DeepSeek V4 Pro](https://noometry.com/compare/deepseek-v4-pro-vs-kimi-k2-5)

## Other Moonshot AI models

-   [Kimi K3](https://noometry.com/models/kimi-k3)59.5
-   [Kimi K2.6](https://noometry.com/models/kimi-k2-6)47.7
-   [Kimi K2 Thinking Turbo](https://noometry.com/models/kimi-k2-thinking-turbo)45.8
-   [Kimi K2.5 Instant](https://noometry.com/models/kimi-k2-5-instant)43.6
-   [Kimi K2.7 Code](https://noometry.com/models/kimi-k2-7-code)43.3
-   [Kimi K2 (Jul 2025)](https://noometry.com/models/kimi-k2)41.2
-   [Kimi K2.7 Code HighSpeed](https://noometry.com/models/kimi-k2-7-code-highspeed)

## Frequently asked questions

### How good is Kimi K2.5?

Kimi K2.5 by Moonshot AI ranks 57th of 354 ranked models on the Noometry Index as of October 2026, with a score of 48.1. Its strongest category is long context, where it ranks 7th. API pricing starts at $0.45 per million input tokens and $2.25 per million output tokens, with a 262K-token context window.

### How much does Kimi K2.5 cost?

Kimi K2.5 costs $0.45 per million input tokens and $2.25 per million output tokens on deepinfra, with cached input at $0.07.

### What is Kimi K2.5's context window?

Kimi K2.5 accepts up to 262K tokens of input and can write up to 262K tokens in one response.

### Is Kimi K2.5 open source?

Yes. Kimi K2.5's weights are downloadable from Hugging Face (moonshotai/Kimi-K2.5); check the license for commercial terms.

### How fast is Kimi K2.5?

Kimi K2.5 generated about 66 output tokens per second in the Kagi LLM Benchmark's timed runs. Speed varies by provider, load and reasoning effort.

### What are Kimi K2.5's strengths and weaknesses?

Relative to other ranked models, Kimi K2.5 places best in long context, coding, math and lowest in agentic & tool use, multimodal, reasoning.

### What is Kimi K2.5 best at?

Its best category is long context, where it ranks 7th on Noometry.

### Cite this page

Noometry. (2026). Kimi K2.5 benchmarks and pricing. Retrieved October 10, 2026, from https://noometry.com/models/kimi-k2-5

Quote Noometry with a link back to this page. It is also available in [Markdown](https://noometry.com/md/models/kimi-k2-5.md).
