Model comparison

Grok 4.6 vs Kimi K3

Kimi K3 is the stronger model overall, scoring 59.5 to 56.9 on the Noometry Index. Grok 4.6 costs 2.0× less per token, which makes it the better buy when Kimi K3's lead doesn't matter for your workload.

Last verified . 45 shared benchmarks.

Grok 4.6 xAI

56.9

Rank #21 Confirmed

Kimi K3 Moonshot AI

59.5

Rank #15 Confirmed

Summary

  • They share 45 benchmarks with published results for both. Grok 4.6 scores higher in 2 categories and Kimi K3 in 8 categories; 9 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Kimi K3 leads 76.6 to 62.3.
  • The biggest single-benchmark swing is ProofBench: 51% for Grok 4.6 and 87% for Kimi K3.
  • Grok 4.6 is cheaper at $2 / $6 per million input/output tokens, against $3 / $15 for Kimi K3.
  • Kimi K3 accepts more context: 1.05M tokens versus 500K.
  • Kimi K3 has downloadable open weights; the other is API-only.

Side by side

Grok 4.6 and Kimi K3 specifications
Grok 4.6Kimi K3
ProviderxAIMoonshot AI
Noometry Index56.959.5
Released2026-08-122026-07-16
WeightsProprietaryOpen
Context window500K1.05M
Max output500K1.05M
Input $ / M tokens$2$3
Output $ / M tokens$6$15
Results tracked4953

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Kimi K3 leads

Grok 4.6: 58.5 (#16), Kimi K3: 61.0 (#10)

Coding benchmarks
BenchmarkGrok 4.6Kimi K3
DeepSWE67.5%68.5%
FrontierCode48%44.2%
LMArena WebDev16171654
FrontierSWE25.3%25.9%
SciCode56.5%59.5%
WeirdML67.3%82.6%
LMArena Coding14651508
ALE-Bench1,5081,524
CursorBench41.4%—

Agentic & Tool Use Kimi K3 leads

Grok 4.6: 39.4 (#27), Kimi K3: 41.8 (#20)

Agentic & Tool Use benchmarks
BenchmarkGrok 4.6Kimi K3
APEX-Agents65.3%50.6%
GDP.pdf17.2%19%
Vending-Bench 29,0475,165
τ²-bench Banking—37.1%
PostTrainBench—32%
GBAEval—48.3%

Reasoning Kimi K3 leads

Grok 4.6: 61.4 (#20), Kimi K3: 63.0 (#17)

Reasoning benchmarks
BenchmarkGrok 4.6Kimi K3
ARC-AGI-267.1%60.4%
SimpleBench75.9%60.7%
NYT Connections (extended)80%93.6%
ARC-AGI-187.5%94.5%
CritPt19.7%23.4%
Chess Puzzles40%39%
LMArena Hard Prompts14471496
Mystery Game Puzzles34%26%
DTBench97.3%91.2%
LMCA48.5%52.7%
Epoch Capabilities Index156.44157.45
EBR-Bench30.5%—
Surface Evolver Bench—95%
ForecastBench—61.1

Math Kimi K3 leads

Grok 4.6: 67.0 (#24), Kimi K3: 74.2 (#16)

Math benchmarks
BenchmarkGrok 4.6Kimi K3
FrontierMath (Tiers 1-3)66%72.2%
FrontierMath Tier 431.7%39%
OTIS Mock AIME 2024-202599.2%97.2%
ProofBench51%87%
LMArena Math14231491
MathArena Final-Answer Competitions—87.8%

Knowledge Too close to call

Grok 4.6: 63.3 (#20), Kimi K3: 63.2 (#21)

Knowledge benchmarks
BenchmarkGrok 4.6Kimi K3
GPQA Diamond94%93.1%
SimpleQA Verified49.3%50.6%
LMArena Expert14671521

Multimodal Grok 4.6 leads

Grok 4.6: 43.6 (#23), Kimi K3: 37.8 (#70)

Multimodal benchmarks
BenchmarkGrok 4.6Kimi K3
Blueprint-Bench 233.2%29.5%
Furniture Assembly40%34.2%
LMArena Vision1263—
LMArena Document1452—

Multilingual Kimi K3 leads

Grok 4.6: 53.0 (#74), Kimi K3: 56.3 (#21)

Multilingual benchmarks
BenchmarkGrok 4.6Kimi K3
LMArena Non-English14201466
LMArena Chinese14801529
LMArena French14611491
LMArena German14311488
LMArena Japanese13761487
LMArena Korean13971458
LMArena Russian14221482
LMArena Spanish14041472

Instruction Following Kimi K3 leads

Grok 4.6: 75.4 (#63), Kimi K3: 77.7 (#14)

Instruction Following benchmarks
BenchmarkGrok 4.6Kimi K3
LMArena Instruction Following14311483

Long Context Kimi K3 leads

Grok 4.6: 44.5 (#66), Kimi K3: 45.8 (#29)

Long Context benchmarks
BenchmarkGrok 4.6Kimi K3
LMArena Longer Query14541494

Writing & Preference Kimi K3 leads

Grok 4.6: 62.3 (#80), Kimi K3: 76.6 (#4)

Writing & Preference benchmarks
BenchmarkGrok 4.6Kimi K3
LMArena Text14281476
LMArena Creative Writing14281454
LMArena Multi-Turn14251488
EQ-Bench Creative Writing—2082
EQ-Bench 4—1339

Frequently asked questions

Is Grok 4.6 better than Kimi K3?

Kimi K3 is the stronger model overall, scoring 59.5 to 56.9 on the Noometry Index. Grok 4.6 costs 2.0× less per token, which makes it the better buy when Kimi K3's lead doesn't matter for your workload.

Which is cheaper, Grok 4.6 or Kimi K3?

Grok 4.6 is cheaper. It lists at $2 per million input tokens and $6 per million output tokens; Kimi K3 lists at $3 and $15.

Is Grok 4.6 or Kimi K3 better for coding?

Kimi K3 scores higher on coding benchmarks: 61.0 versus 58.5 in the Noometry coding category.

Which has the bigger context window?

Kimi K3 does, with 1.05M tokens against 500K.

How many benchmarks do Grok 4.6 and Kimi K3 share?

45 benchmarks have published results for both models. Grok 4.6 has 49 scored results on Noometry and Kimi K3 has 53.

Related comparisons

Go deeper