Moonshot AI, open weights
Kimi K2 (Jul 2025)
Kimi K2 (Jul 2025) by Moonshot AI ranks 140th of 354 ranked models on the Noometry Index as of October 2026, with a score of 41.2. Its strongest category is agentic & tool use, where it ranks 64th. API pricing starts at $0.57 per million input tokens and $2.30 per million output tokens, with a 262K-token context window.
Last verified
Specifications
- Noometry rank
- #140 of 354
- Index score
- 41.2
- Evidence
- Confirmed 42 results
- Provider
- Moonshot AI
- Released
- July 12, 2025
- Weights
- Open weights
- Reasoning
- Yes
- Context window
- 262K
- Max output
- 262K
- Input price
- $0.57 / M
- Output price
- $2.30 / M
- Blended price
- $1 / M
- Output speed
- 201 tokens/s Kagi
- Value
- #116 of 219
- Knowledge cutoff
- August 2024
- Input
- text
- Hugging Face
- moonshotai/Kimi-K2-Instruct
Category scores
Each category score combines every public result we have in that category.
- Coding 42.4
- Agentic & Tool Use 32.4
- Reasoning 23.3
- Math 42.7
- Knowledge 37.3
- Multilingual 49.6
- Instruction Following 71.1
- Long Context 41.2
- Writing & Preference 62.3
| Category | Score | Rank | Results |
|---|---|---|---|
| Coding | 42.4 | #102 | 5 |
| Agentic & Tool Use | 32.4 | #64 | 2 |
| Reasoning | 23.3 | #179 | 3 |
| Math | 42.7 | #83 | 2 |
| Knowledge | 37.3 | #157 | 5 |
| Multilingual | 49.6 | #130 | 1 |
| Instruction Following | 71.1 | #156 | 2 |
| Long Context | 41.2 | #145 | 3 |
| Writing & Preference | 62.3 | #78 | 6 |
Strengths and weaknesses
Categories where Kimi K2 (Jul 2025) places highest and lowest among the models ranked in each, with its score against that category's median.
Strongest categories
| Category | Score | vs median | Rank |
|---|---|---|---|
| Writing & Preference | 62.3 | +8.5 | #78 of 312, top 25% |
| Math | 42.7 | +6.1 | #83 of 327, top 26% |
| Coding | 42.4 | +3.7 | #102 of 340, top 30% |
Weakest categories
| Category | Score | vs median | Rank |
|---|---|---|---|
| Instruction Following | 71.1 | −0.1 | #156 of 305, top 52% |
| Reasoning | 23.3 | −0.3 | #179 of 350, top 52% |
| Knowledge | 37.3 | −0.0 | #157 of 314, top 50% |
Closest competitors
The models ranked just above and below Kimi K2 (Jul 2025). When scores are this close, price and speed are often the better way to choose.
| Model | Rank | Score | Blended $/M | Speed | |
|---|---|---|---|---|---|
| Grok 4.1 Fast | #136 | 41.4 | $0.28 | — | Compare |
| GLM-4.6V | #137 | 41.3 | $0.45 | — | Compare |
| MiMo-V2-Flash | #138 | 41.3 | $0.18 | — | Compare |
| Hunyuan Turbos 20250226 | #139 | 41.3 | — | — | Compare |
| Grok-3 mini | #141 | 41.2 | — | 10 | Compare |
| Claude Opus 4.1 | #142 | 41.0 | $30 | — | Compare |
| o1 | #143 | 40.9 | $26.25 | — | Compare |
| Gemini 3.1 Flash Lite | #144 | 40.8 | $0.56 | 10 | Compare |
Sponsored placements are available on pages like this one. Advertise on Noometry
Benchmark results
Every published result we track, with its source. Bold rows are the ones used for ranking; where several exist we prefer independent runs over self-reported numbers.
Coding
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| SWE-bench Verified (bash only) | 43.8% | SWE-bench | 2025-08-07 | ||
| SWE-bench Verified (bash only) | 63.4% | #18 of 39, top 47% | SWE-bench | 2025-12-10 | |
| Aider Polyglot | 59.1% | #16 of 44, top 37% | Epoch AI | ||
| Aider Polyglot | 59.1% | #16 of 44, top 37% | Epoch AI | ||
| GSO | 4.9% | #22 of 31, top 71% | Epoch AI | ||
| WeirdML | 39.4% | Epoch AI | |||
| WeirdML | 36.7% | Epoch AI | |||
| WeirdML | 42.8% | #68 of 119, top 58% | Epoch AI | ||
| WeirdML | 39.4% | Epoch AI | |||
| LMArena Coding | 1399 | #137 of 294, top 47% | LMArena | 2026-10-08 | |
| LMArena Coding | 1377 | LMArena | 2026-10-08 | ||
| ALE-Bench | 597.5 | #81 of 105, top 78% | Epoch AI | ||
| ALE-Bench | 267.12 | Epoch AI |
Agentic & Tool Use
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| Terminal-Bench | 27.8% | Epoch AI | |||
| Terminal-Bench | 35.7% | #27 of 41, top 66% | Epoch AI | ||
| Berkeley Function Calling Leaderboard | 59.1% | #9 of 49, top 19% | fc | Berkeley Function Calling Leaderboard | |
| METR Time Horizons | 59.2% | #19 of 32, top 60% | Epoch AI |
Reasoning
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| SimpleBench | 26.3% | #65 of 77, top 85% | Epoch AI | ||
| Kagi LLM Benchmark | 64.4% | #33 of 99, top 34% | Kagi LLM Benchmark | ||
| Kagi LLM Benchmark | 52.2% | Kagi LLM Benchmark | |||
| Kagi LLM Benchmark | 45% | Kagi LLM Benchmark | |||
| LMArena Hard Prompts | 1366 | LMArena | 2026-10-08 | ||
| LMArena Hard Prompts | 1384 | #136 of 297, top 46% | LMArena | 2026-10-08 | |
| Epoch Capabilities Index | 146.01 | #76 of 213, top 36% | Epoch AI | 2025-11-06 | |
| Epoch Capabilities Index | 140.13 | Epoch AI | 2025-07-12 | ||
| ForecastBench | 60.2 | #30 of 72, top 42% | Epoch AI | ||
| ForecastBench | 59.8 | Epoch AI |
Math
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| Omni-MATH | 65.4% | #6 of 57, top 11% | HELM Capabilities | ||
| LMArena Math | 1397 | #127 of 285, top 45% | LMArena | 2026-10-08 | |
| LMArena Math | 1367 | LMArena | 2026-10-08 | ||
| FrontierMath (Feb 2025 set) | 21.4% | #28 of 68, top 42% | Epoch AI | 2025-12-05 | |
| FrontierMath Tier 4 (v1) | 0% | #52 of 55, top 95% | Epoch AI | 2025-12-04 |
Knowledge
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| MMLU-Pro | 81.9% | #13 of 58, top 23% | HELM Capabilities | ||
| Confabulations (lower is better) | 20.4% | #34 of 51, top 67% | Lech Mazur benchmarks | ||
| Vectara Hallucination Rate (lower is better) | 17.9% | #88 of 96, top 92% | Vectara Hallucination Leaderboard | ||
| GPQA (HELM) | 65.3% | #17 of 57, top 30% | HELM Capabilities | ||
| LMArena Expert | 1365 | #143 of 273, top 53% | LMArena | 2026-10-08 | |
| LMArena Expert | 1346 | LMArena | 2026-10-08 |
Multilingual
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Non-English | 1372 | #130 of 297, top 44% | LMArena | 2026-10-08 | |
| LMArena Non-English | 1358 | LMArena | 2026-10-08 | ||
| LMArena Chinese | 1415 | #131 of 285, top 46% | LMArena | 2026-10-08 | |
| LMArena Chinese | 1401 | LMArena | 2026-10-08 | ||
| LMArena French | 1373 | LMArena | 2026-10-08 | ||
| LMArena French | 1379 | #132 of 223, top 60% | LMArena | 2026-10-08 | |
| LMArena German | 1387 | #106 of 231, top 46% | LMArena | 2026-10-08 | |
| LMArena German | 1374 | LMArena | 2026-10-08 | ||
| LMArena Japanese | 1336 | LMArena | 2026-10-08 | ||
| LMArena Japanese | 1349 | #98 of 211, top 47% | LMArena | 2026-10-08 | |
| LMArena Korean | 1325 | #116 of 213, top 55% | LMArena | 2026-10-08 | |
| LMArena Korean | 1287 | LMArena | 2026-10-08 | ||
| LMArena Russian | 1361 | LMArena | 2026-10-08 | ||
| LMArena Russian | 1385 | #126 of 283, top 45% | LMArena | 2026-10-08 | |
| LMArena Spanish | 1386 | #122 of 226, top 54% | LMArena | 2026-10-08 | |
| LMArena Spanish | 1330 | LMArena | 2026-10-08 |
Instruction Following
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| IFEval | 85% | #18 of 57, top 32% | HELM Capabilities | ||
| LMArena Instruction Following | 1324 | LMArena | 2026-10-08 | ||
| LMArena Instruction Following | 1348 | #146 of 298, top 49% | LMArena | 2026-10-08 |
Long Context
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| Fiction.LiveBench | 61.1% | Epoch AI | |||
| Fiction.LiveBench | 66.7% | #21 of 47, top 45% | Epoch AI | ||
| CL-bench | 11.9% | Epoch AI | |||
| CL-bench | 17.6% | #13 of 19, top 69% | Epoch AI | ||
| LMArena Longer Query | 1326 | LMArena | 2026-10-08 | ||
| LMArena Longer Query | 1353 | #151 of 291, top 52% | LMArena | 2026-10-08 |
Writing & Preference
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Text | 1380 | #135 of 297, top 46% | LMArena | 2026-10-08 | |
| LMArena Text | 1371 | LMArena | 2026-10-08 | ||
| LMArena Creative Writing | 1324 | LMArena | 2026-10-08 | ||
| LMArena Creative Writing | 1350 | #127 of 295, top 44% | LMArena | 2026-10-08 | |
| Short-Story Creative Writing | 85.6% | #2 of 39, top 6% | Epoch AI | ||
| EQ-Bench Creative Writing | 1631 | EQ-Bench | |||
| EQ-Bench Creative Writing | 1666 | #37 of 115, top 33% | EQ-Bench | ||
| WildBench | 86.2% | #3 of 57, top 6% | HELM Capabilities | ||
| LMArena Multi-Turn | 1365 | LMArena | 2026-10-08 | ||
| LMArena Multi-Turn | 1371 | #139 of 295, top 48% | LMArena | 2026-10-08 |
API pricing by provider
| Route | Input $/M | Output $/M | Cached input $/M | Checked |
|---|---|---|---|---|
| bedrock | $0.60 | $2.50 | — | 2026-10-10 |
| openrouter | $0.57 | $2.30 | — | 2026-10-10 |
| vertex | $0.60 | $2.50 | $0.06 | 2026-10-10 |
Compare Kimi K2 (Jul 2025)
- Kimi K2 (Jul 2025) vs Hunyuan Turbos 20250226
- Kimi K2 (Jul 2025) vs Grok-3 mini
- Kimi K2 (Jul 2025) vs MiMo-V2-Flash
- Kimi K2 (Jul 2025) vs Claude Opus 4.1
- Kimi K2 (Jul 2025) vs GLM-4.6V
- Kimi K2 (Jul 2025) vs o1
- Kimi K2 (Jul 2025) vs GPT-6 Astra
- Kimi K2 (Jul 2025) vs Claude Fable 5.1
- Kimi K2 (Jul 2025) vs Gemini 3.8 Flash
- Kimi K2 (Jul 2025) vs Grok 4.6
- Kimi K2 (Jul 2025) vs Qwen3.8 Max
- Kimi K2 (Jul 2025) vs GLM-5.3
- Kimi K2 (Jul 2025) vs Muse Spark 1.3
- Kimi K2 (Jul 2025) vs DeepSeek V4 Pro
Other Moonshot AI models
Frequently asked questions
How good is Kimi K2 (Jul 2025)?
Kimi K2 (Jul 2025) by Moonshot AI ranks 140th of 354 ranked models on the Noometry Index as of October 2026, with a score of 41.2. Its strongest category is agentic & tool use, where it ranks 64th. API pricing starts at $0.57 per million input tokens and $2.30 per million output tokens, with a 262K-token context window.
How much does Kimi K2 (Jul 2025) cost?
Kimi K2 (Jul 2025) costs $0.57 per million input tokens and $2.30 per million output tokens on openrouter.
What is Kimi K2 (Jul 2025)'s context window?
Kimi K2 (Jul 2025) accepts up to 262K tokens of input and can write up to 262K tokens in one response.
Is Kimi K2 (Jul 2025) open source?
Yes. Kimi K2 (Jul 2025)'s weights are downloadable from Hugging Face (moonshotai/Kimi-K2-Instruct); check the license for commercial terms.
How fast is Kimi K2 (Jul 2025)?
Kimi K2 (Jul 2025) generated about 201 output tokens per second in the Kagi LLM Benchmark's timed runs. Speed varies by provider, load and reasoning effort.
What are Kimi K2 (Jul 2025)'s strengths and weaknesses?
Relative to other ranked models, Kimi K2 (Jul 2025) places best in writing & preference, math, coding and lowest in instruction following, reasoning, knowledge.
What is Kimi K2 (Jul 2025) best at?
Its best category is agentic & tool use, where it ranks 64th on Noometry.