Moonshot AI, open weights
Kimi K3
Kimi K3 by Moonshot AI ranks 15th of 354 ranked models on the Noometry Index as of October 2026, with a score of 59.5. Its strongest category is writing & preference, where it ranks 4th. API pricing starts at $3 per million input tokens and $15 per million output tokens, with a 1.05M-token context window.
Last verified
Specifications
- Noometry rank
- #15 of 354
- Index score
- 59.5
- Evidence
- Confirmed 53 results
- Provider
- Moonshot AI
- Released
- July 16, 2026
- Weights
- Open weights
- Reasoning
- Yes
- Context window
- 1.05M
- Max output
- 1.05M
- Input price
- $3 / M
- Output price
- $15 / M
- Blended price
- $6 / M
- Output speed
- Not measured
- Value
- #185 of 219
- Knowledge cutoff
- Unknown
- Input
- text, image, video
- Hugging Face
- moonshotai/Kimi-K3
Category scores
Each category score combines every public result we have in that category.
- Coding 61.0
- Agentic & Tool Use 41.8
- Reasoning 63.0
- Math 74.2
- Knowledge 63.2
- Multimodal 37.8
- Multilingual 56.3
- Instruction Following 77.7
- Long Context 45.8
- Writing & Preference 76.6
| Category | Score | Rank | Results |
|---|---|---|---|
| Coding | 61.0 | #10 | 7 |
| Agentic & Tool Use | 41.8 | #20 | 5 |
| Reasoning | 63.0 | #17 | 11 |
| Math | 74.2 | #16 | 6 |
| Knowledge | 63.2 | #21 | 3 |
| Multimodal | 37.8 | #70 | 2 |
| Multilingual | 56.3 | #21 | 1 |
| Instruction Following | 77.7 | #14 | 1 |
| Long Context | 45.8 | #29 | 1 |
| Writing & Preference | 76.6 | #4 | 5 |
Strengths and weaknesses
Categories where Kimi K3 places highest and lowest among the models ranked in each, with its score against that category's median.
Strongest categories
| Category | Score | vs median | Rank |
|---|---|---|---|
| Writing & Preference | 76.6 | +22.8 | #4 of 312, top 2% |
| Coding | 61.0 | +22.3 | #10 of 340, top 3% |
| Instruction Following | 77.7 | +6.5 | #14 of 305, top 5% |
Weakest categories
| Category | Score | vs median | Rank |
|---|---|---|---|
| Multimodal | 37.8 | −0.8 | #70 of 128, top 55% |
| Agentic & Tool Use | 41.8 | +11.5 | #20 of 154, top 13% |
| Long Context | 45.8 | +4.9 | #29 of 296, top 10% |
Closest competitors
The models ranked just above and below Kimi K3. When scores are this close, price and speed are often the better way to choose.
| Model | Rank | Score | Blended $/M | Speed | |
|---|---|---|---|---|---|
| Gemini 3.8 Flash | #11 | 61.8 | $1.50 | — | Compare |
| GPT-6 Sol | #12 | 61.8 | $4 | — | Compare |
| Claude Opus 4.8 | #13 | 60.7 | $10 | 34 | Compare |
| Gemini 3.7 Flash | #14 | 59.8 | $1.50 | — | Compare |
| GPT-5.4 | #16 | 59.4 | $5.63 | 12 | Compare |
| GPT-5.6 Terra | #17 | 59.2 | $4.50 | 11 | Compare |
| GPT-5.4 Pro | #18 | 58.9 | $67.50 | — | Compare |
| Claude Opus 4.7 | #19 | 58.3 | $10 | 33 | Compare |
Sponsored placements are available on pages like this one. Advertise on Noometry
Benchmark results
Every published result we track, with its source. Bold rows are the ones used for ranking; where several exist we prefer independent runs over self-reported numbers.
Coding
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| DeepSWE | 68.5% | #10 of 29, top 35% | max | Epoch AI | |
| FrontierCode | 44.2% | #13 of 37, top 36% | Epoch AI | ||
| LMArena WebDev | 1654 | #11 of 113, top 10% | LMArena | 2026-10-08 | |
| FrontierSWE | 25.9% | #11 of 18, top 62% | max | Epoch AI | |
| SciCode | 58.7% | Epoch AI | |||
| SciCode | 51.2% | low | Epoch AI | ||
| SciCode | 59.5% | #9 of 121, top 8% | max | Epoch AI | |
| WeirdML | 82.6% | #9 of 119, top 8% | max | Epoch AI | |
| LMArena Coding | 1508 | #12 of 294, top 5% | LMArena | 2026-10-08 | |
| ALE-Bench | 1,524 | #16 of 105, top 16% | max | Epoch AI |
Agentic & Tool Use
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| APEX-Agents | 50.6% | #25 of 49, top 52% | Epoch AI | ||
| τ²-bench Banking | 37.1% | #12 of 26, top 47% | max | τ²-bench | 2026-08-04 |
| PostTrainBench | 32% | #5 of 11, top 46% | Epoch AI | ||
| GBAEval | 48.3% | #9 of 23, top 40% | Epoch AI | ||
| GDP.pdf | 19% | #21 of 36, top 59% | max | Epoch AI | |
| Vending-Bench 2 | 5,165 | #27 of 60, top 45% | Epoch AI |
Reasoning
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| ARC-AGI-2 | 55% | high | Epoch AI | ||
| ARC-AGI-2 | 12.4% | low | Epoch AI | ||
| ARC-AGI-2 | 60.4% | #29 of 83, top 35% | max | Epoch AI | |
| SimpleBench | 60.7% | #23 of 77, top 30% | max | Epoch AI | |
| NYT Connections (extended) | 93.6% | #10 of 91, top 11% | Lech Mazur benchmarks | ||
| ARC-AGI-1 | 86.7% | high | Epoch AI | ||
| ARC-AGI-1 | 65.7% | low | Epoch AI | ||
| ARC-AGI-1 | 94.5% | #17 of 83, top 21% | max | Epoch AI | |
| CritPt | 3.1% | low | Epoch AI | ||
| CritPt | 23.4% | #19 of 134, top 15% | max | Epoch AI | |
| Chess Puzzles | 25% | high | Epoch AI | 2026-08-07 | |
| Chess Puzzles | 20% | low | Epoch AI | 2026-08-07 | |
| Chess Puzzles | 39% | #23 of 129, top 18% | max | Epoch AI | 2026-07-16 |
| LMArena Hard Prompts | 1496 | #12 of 297, top 5% | LMArena | 2026-10-08 | |
| Mystery Game Puzzles | 12% | low | Epoch AI | 2026-08-29 | |
| Mystery Game Puzzles | 26% | #32 of 74, top 44% | max | Epoch AI | 2026-07-29 |
| DTBench | 91.2% | #31 of 151, top 21% | max | Epoch AI | |
| LMCA | 52.7% | #17 of 125, top 14% | max | Epoch AI | |
| Surface Evolver Bench | 95% | #2 of 25, top 8% | Epoch AI | ||
| Surface Evolver Bench | 93% | max | Epoch AI | ||
| Epoch Capabilities Index | 157.45 | #15 of 213, top 8% | Epoch AI | 2026-07-16 | |
| ForecastBench | 61.1 | #17 of 72, top 24% | max | Epoch AI |
Math
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| FrontierMath (Tiers 1-3) | 72.2% | #22 of 81, top 28% | max | Epoch AI | 2026-07-17 |
| FrontierMath Tier 4 | 39% | #23 of 63, top 37% | max | Epoch AI | 2026-07-17 |
| MathArena Final-Answer Competitions | 87.8% | #3 of 29, top 11% | think | MathArena | |
| OTIS Mock AIME 2024-2025 | 93.3% | high | Epoch AI | 2026-08-07 | |
| OTIS Mock AIME 2024-2025 | 68.9% | low | Epoch AI | 2026-08-07 | |
| OTIS Mock AIME 2024-2025 | 97.2% | #28 of 173, top 17% | max | Epoch AI | 2026-07-16 |
| ProofBench | 87% | #9 of 77, top 12% | Epoch AI | ||
| LMArena Math | 1491 | #17 of 285, top 6% | LMArena | 2026-10-08 |
Knowledge
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| GPQA Diamond | 91.9% | high | Epoch AI | 2026-08-07 | |
| GPQA Diamond | 84.8% | low | Epoch AI | 2026-08-07 | |
| GPQA Diamond | 93.1% | #18 of 186, top 10% | max | Epoch AI | 2026-07-16 |
| SimpleQA Verified | 50.6% | #24 of 77, top 32% | max | Epoch AI | 2026-08-27 |
| LMArena Expert | 1521 | #10 of 273, top 4% | LMArena | 2026-10-08 |
Multimodal
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| Blueprint-Bench 2 | 29.5% | #17 of 31, top 55% | Epoch AI | ||
| Furniture Assembly | 34.2% | #20 of 31, top 65% | max | Epoch AI | 2026-09-11 |
Multilingual
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Non-English | 1466 | #21 of 297, top 8% | LMArena | 2026-10-08 | |
| LMArena Chinese | 1529 | #15 of 285, top 6% | LMArena | 2026-10-08 | |
| LMArena French | 1491 | #19 of 223, top 9% | LMArena | 2026-10-08 | |
| LMArena German | 1488 | #14 of 231, top 7% | LMArena | 2026-10-08 | |
| LMArena Japanese | 1487 | #11 of 211, top 6% | LMArena | 2026-10-08 | |
| LMArena Korean | 1458 | #14 of 213, top 7% | LMArena | 2026-10-08 | |
| LMArena Russian | 1482 | #18 of 283, top 7% | LMArena | 2026-10-08 | |
| LMArena Spanish | 1472 | #20 of 226, top 9% | LMArena | 2026-10-08 |
Instruction Following
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Instruction Following | 1483 | #11 of 298, top 4% | LMArena | 2026-10-08 |
Long Context
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Longer Query | 1494 | #11 of 291, top 4% | LMArena | 2026-10-08 |
Writing & Preference
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Text | 1476 | #20 of 297, top 7% | LMArena | 2026-10-08 | |
| LMArena Creative Writing | 1454 | #25 of 295, top 9% | LMArena | 2026-10-08 | |
| EQ-Bench Creative Writing | 2082 | #5 of 115, top 5% | EQ-Bench | ||
| EQ-Bench 4 | 1339 | #3 of 28, top 11% | EQ-Bench | ||
| LMArena Multi-Turn | 1488 | #13 of 295, top 5% | LMArena | 2026-10-08 |
API pricing by provider
Compare Kimi K3
- Kimi K3 vs Kimi K2.6
- Kimi K3 vs Gemini 3.7 Flash
- Kimi K3 vs GPT-5.4
- Kimi K3 vs Claude Opus 4.8
- Kimi K3 vs GPT-5.6 Terra
- Kimi K3 vs GPT-6 Sol
- Kimi K3 vs GPT-5.4 Pro
- Kimi K3 vs GPT-6 Astra
- Kimi K3 vs Claude Fable 5.1
- Kimi K3 vs Gemini 3.8 Flash
- Kimi K3 vs Grok 4.6
- Kimi K3 vs Qwen3.8 Max
- Kimi K3 vs GLM-5.3
- Kimi K3 vs Muse Spark 1.3
Other Moonshot AI models
Frequently asked questions
How good is Kimi K3?
Kimi K3 by Moonshot AI ranks 15th of 354 ranked models on the Noometry Index as of October 2026, with a score of 59.5. Its strongest category is writing & preference, where it ranks 4th. API pricing starts at $3 per million input tokens and $15 per million output tokens, with a 1.05M-token context window.
How much does Kimi K3 cost?
Kimi K3 costs $3 per million input tokens and $15 per million output tokens on Moonshot AI's own API, with cached input at $0.30.
What is Kimi K3's context window?
Kimi K3 accepts up to 1.05M tokens of input and can write up to 1.05M tokens in one response.
Is Kimi K3 open source?
Yes. Kimi K3's weights are downloadable from Hugging Face (moonshotai/Kimi-K3); check the license for commercial terms.
What are Kimi K3's strengths and weaknesses?
Relative to other ranked models, Kimi K3 places best in writing & preference, coding, instruction following and lowest in multimodal, agentic & tool use, long context.
What is Kimi K3 best at?
Its best category is writing & preference, where it ranks 4th on Noometry.