Model comparison

Grok 4.5 vs Kimi K3

Kimi K3 is the stronger model overall, scoring 59.5 to 55.0 on the Noometry Index. Grok 4.5 costs 2.0× less per token, which makes it the better buy when Kimi K3's lead doesn't matter for your workload.

Last verified . 48 shared benchmarks.

Grok 4.5 xAI

55.0

Rank #25 Confirmed

Kimi K3 Moonshot AI

59.5

Rank #15 Confirmed

Summary

  • They share 48 benchmarks with published results for both. Grok 4.5 scores higher in 1 category and Kimi K3 in 9 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in math, where Kimi K3 leads 74.2 to 60.9.
  • The biggest single-benchmark swing is ProofBench: 31% for Grok 4.5 and 87% for Kimi K3.
  • Grok 4.5 is cheaper at $2 / $6 per million input/output tokens, against $3 / $15 for Kimi K3.
  • Kimi K3 accepts more context: 1.05M tokens versus 500K.
  • Kimi K3 has downloadable open weights; the other is API-only.

Side by side

Grok 4.5 and Kimi K3 specifications
Grok 4.5Kimi K3
ProviderxAIMoonshot AI
Noometry Index55.059.5
Released2026-07-082026-07-16
WeightsProprietaryOpen
Context window500K1.05M
Max output500K1.05M
Input $ / M tokens$2$3
Output $ / M tokens$6$15
Results tracked5253

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Kimi K3 leads

Grok 4.5: 52.2 (#35), Kimi K3: 61.0 (#10)

Coding benchmarks
BenchmarkGrok 4.5Kimi K3
DeepSWE53.8%68.5%
FrontierCode42.4%44.2%
LMArena WebDev15531654
SciCode54.1%59.5%
WeirdML46.4%82.6%
LMArena Coding14741508
ALE-Bench1,3091,524
FrontierSWE—25.9%

Agentic & Tool Use Grok 4.5 leads

Grok 4.5: 44.4 (#17), Kimi K3: 41.8 (#20)

Agentic & Tool Use benchmarks
BenchmarkGrok 4.5Kimi K3
APEX-Agents56.2%50.6%
τ²-bench Banking47.9%37.1%
PostTrainBench23.4%32%
GBAEval65.4%48.3%
GDP.pdf14%19%
Vending-Bench 23,8875,165
LMArena Search1213—

Reasoning Kimi K3 leads

Grok 4.5: 56.1 (#25), Kimi K3: 63.0 (#17)

Reasoning benchmarks
BenchmarkGrok 4.5Kimi K3
ARC-AGI-252.6%60.4%
SimpleBench70%60.7%
NYT Connections (extended)79.9%93.6%
ARC-AGI-187.2%94.5%
CritPt15.4%23.4%
Chess Puzzles36%39%
LMArena Hard Prompts14621496
DTBench96.5%91.2%
LMCA45.2%52.7%
Surface Evolver Bench74.4%95%
Epoch Capabilities Index153.92157.45
Kagi LLM Benchmark83.5%—
Mystery Game Puzzles—26%
ForecastBench—61.1

Math Kimi K3 leads

Grok 4.5: 60.9 (#35), Kimi K3: 74.2 (#16)

Math benchmarks
BenchmarkGrok 4.5Kimi K3
FrontierMath (Tiers 1-3)57.2%72.2%
FrontierMath Tier 424.4%39%
OTIS Mock AIME 2024-202597.8%97.2%
ProofBench31%87%
LMArena Math14591491
MathArena Final-Answer Competitions—87.8%

Knowledge Too close to call

Grok 4.5: 62.3 (#24), Kimi K3: 63.2 (#21)

Knowledge benchmarks
BenchmarkGrok 4.5Kimi K3
GPQA Diamond93.4%93.1%
SimpleQA Verified48.3%50.6%
LMArena Expert14661521

Multimodal Too close to call

Grok 4.5: 37.6 (#72), Kimi K3: 37.8 (#70)

Multimodal benchmarks
BenchmarkGrok 4.5Kimi K3
Blueprint-Bench 227.3%29.5%
Furniture Assembly22.5%34.2%
LMArena Vision1288—
LMArena Document1452—

Multilingual Kimi K3 leads

Grok 4.5: 54.4 (#42), Kimi K3: 56.3 (#21)

Multilingual benchmarks
BenchmarkGrok 4.5Kimi K3
LMArena Non-English14401466
LMArena Chinese14961529
LMArena French14561491
LMArena German14461488
LMArena Japanese14281487
LMArena Korean14041458
LMArena Russian14481482
LMArena Spanish14501472

Instruction Following Kimi K3 leads

Grok 4.5: 76.0 (#48), Kimi K3: 77.7 (#14)

Instruction Following benchmarks
BenchmarkGrok 4.5Kimi K3
LMArena Instruction Following14461483

Long Context Kimi K3 leads

Grok 4.5: 44.8 (#56), Kimi K3: 45.8 (#29)

Long Context benchmarks
BenchmarkGrok 4.5Kimi K3
LMArena Longer Query14631494

Writing & Preference Kimi K3 leads

Grok 4.5: 65.8 (#42), Kimi K3: 76.6 (#4)

Writing & Preference benchmarks
BenchmarkGrok 4.5Kimi K3
LMArena Text14481476
LMArena Creative Writing14421454
EQ-Bench Creative Writing15792082
LMArena Multi-Turn14561488
EQ-Bench 4—1339

Frequently asked questions

Is Grok 4.5 better than Kimi K3?

Kimi K3 is the stronger model overall, scoring 59.5 to 55.0 on the Noometry Index. Grok 4.5 costs 2.0× less per token, which makes it the better buy when Kimi K3's lead doesn't matter for your workload.

Which is cheaper, Grok 4.5 or Kimi K3?

Grok 4.5 is cheaper. It lists at $2 per million input tokens and $6 per million output tokens; Kimi K3 lists at $3 and $15.

Is Grok 4.5 or Kimi K3 better for coding?

Kimi K3 scores higher on coding benchmarks: 61.0 versus 52.2 in the Noometry coding category.

Which has the bigger context window?

Kimi K3 does, with 1.05M tokens against 500K.

How many benchmarks do Grok 4.5 and Kimi K3 share?

48 benchmarks have published results for both models. Grok 4.5 has 52 scored results on Noometry and Kimi K3 has 53.

Related comparisons

Go deeper