Model comparison

Grok 4.7 vs Kimi K3

Kimi K3 is the stronger model overall, scoring 59.5 to 53.1 on the Noometry Index. Grok 4.7 costs 2.0× less per token, which makes it the better buy when Kimi K3's lead doesn't matter for your workload.

Last verified . 37 shared benchmarks.

Grok 4.7 xAI

53.1

Rank #37 Confirmed

Kimi K3 Moonshot AI

59.5

Rank #15 Confirmed

Summary

  • They share 37 benchmarks with published results for both. Grok 4.7 scores higher in 0 categories and Kimi K3 in 10 categories; 9 gaps are clear of the uncertainty.
  • The widest gap is in math, where Kimi K3 leads 74.2 to 57.8.
  • The biggest single-benchmark swing is ProofBench: 34% for Grok 4.7 and 87% for Kimi K3.
  • Grok 4.7 is cheaper at $2 / $6 per million input/output tokens, against $3 / $15 for Kimi K3.
  • Kimi K3 accepts more context: 1.05M tokens versus 500K.
  • Kimi K3 has downloadable open weights; the other is API-only.

Side by side

Grok 4.7 and Kimi K3 specifications
Grok 4.7Kimi K3
ProviderxAIMoonshot AI
Noometry Index53.159.5
Released2026-09-212026-07-16
WeightsProprietaryOpen
Context window500K1.05M
Max output500K1.05M
Input $ / M tokens$2$3
Output $ / M tokens$6$15
Results tracked3953

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Kimi K3 leads

Grok 4.7: 58.0 (#18), Kimi K3: 61.0 (#10)

Coding benchmarks
BenchmarkGrok 4.7Kimi K3
FrontierCode47.6%44.2%
LMArena WebDev16391654
FrontierSWE29.5%25.9%
SciCode57.8%59.5%
LMArena Coding14271508
DeepSWE—68.5%
CursorBench46.3%—
WeirdML—82.6%
ALE-Bench—1,524

Agentic & Tool Use Kimi K3 leads

Grok 4.7: 36.7 (#37), Kimi K3: 41.8 (#20)

Agentic & Tool Use benchmarks
BenchmarkGrok 4.7Kimi K3
APEX-Agents54.6%50.6%
GDP.pdf22.8%19%
Vending-Bench 210,5375,165
τ²-bench Banking—37.1%
PostTrainBench—32%
GBAEval—48.3%

Reasoning Kimi K3 leads

Grok 4.7: 49.1 (#40), Kimi K3: 63.0 (#17)

Reasoning benchmarks
BenchmarkGrok 4.7Kimi K3
NYT Connections (extended)76.8%93.6%
CritPt18%23.4%
Chess Puzzles38%39%
LMArena Hard Prompts14131496
Mystery Game Puzzles29%26%
DTBench96%91.2%
LMCA49.4%52.7%
Epoch Capabilities Index153.53157.45
ARC-AGI-2—60.4%
SimpleBench—60.7%
ARC-AGI-1—94.5%
Surface Evolver Bench—95%
ForecastBench—61.1

Math Kimi K3 leads

Grok 4.7: 57.8 (#39), Kimi K3: 74.2 (#16)

Math benchmarks
BenchmarkGrok 4.7Kimi K3
FrontierMath (Tiers 1-3)53%72.2%
FrontierMath Tier 417.1%39%
OTIS Mock AIME 2024-202598.1%97.2%
ProofBench34%87%
LMArena Math14071491
MathArena Final-Answer Competitions—87.8%

Knowledge Too close to call

Grok 4.7: 62.8 (#22), Kimi K3: 63.2 (#21)

Knowledge benchmarks
BenchmarkGrok 4.7Kimi K3
GPQA Diamond92.7%93.1%
SimpleQA Verified56%50.6%
LMArena Expert14221521

Multimodal Kimi K3 leads

Grok 4.7: 35.5 (#87), Kimi K3: 37.8 (#70)

Multimodal benchmarks
BenchmarkGrok 4.7Kimi K3
Blueprint-Bench 232.5%29.5%
Furniture Assembly20.8%34.2%
LMArena Vision1228—

Multilingual Kimi K3 leads

Grok 4.7: 50.8 (#116), Kimi K3: 56.3 (#21)

Multilingual benchmarks
BenchmarkGrok 4.7Kimi K3
LMArena Non-English13891466
LMArena Chinese14551529
LMArena French14551491
LMArena Russian13971482
LMArena Spanish14001472
LMArena German—1488
LMArena Japanese—1487
LMArena Korean—1458

Instruction Following Kimi K3 leads

Grok 4.7: 74.1 (#105), Kimi K3: 77.7 (#14)

Instruction Following benchmarks
BenchmarkGrok 4.7Kimi K3
LMArena Instruction Following14041483

Long Context Kimi K3 leads

Grok 4.7: 43.1 (#104), Kimi K3: 45.8 (#29)

Long Context benchmarks
BenchmarkGrok 4.7Kimi K3
LMArena Longer Query14131494

Writing & Preference Kimi K3 leads

Grok 4.7: 70.0 (#24), Kimi K3: 76.6 (#4)

Writing & Preference benchmarks
BenchmarkGrok 4.7Kimi K3
LMArena Text13991476
LMArena Creative Writing13911454
EQ-Bench Creative Writing20072082
LMArena Multi-Turn13931488
EQ-Bench 4—1339

Frequently asked questions

Is Grok 4.7 better than Kimi K3?

Kimi K3 is the stronger model overall, scoring 59.5 to 53.1 on the Noometry Index. Grok 4.7 costs 2.0× less per token, which makes it the better buy when Kimi K3's lead doesn't matter for your workload.

Which is cheaper, Grok 4.7 or Kimi K3?

Grok 4.7 is cheaper. It lists at $2 per million input tokens and $6 per million output tokens; Kimi K3 lists at $3 and $15.

Is Grok 4.7 or Kimi K3 better for coding?

Kimi K3 scores higher on coding benchmarks: 61.0 versus 58.0 in the Noometry coding category.

Which has the bigger context window?

Kimi K3 does, with 1.05M tokens against 500K.

How many benchmarks do Grok 4.7 and Kimi K3 share?

37 benchmarks have published results for both models. Grok 4.7 has 39 scored results on Noometry and Kimi K3 has 53.

Related comparisons

Go deeper