Model comparison

Grok 4.3 vs Kimi K2.5 Instant

Grok 4.3 and Kimi K2.5 Instant score almost the same on the Noometry Index (43.8 vs 43.6), so choose on price, context window or the category you care about most.

Last verified . 18 shared benchmarks.

Grok 4.3 xAI

43.8

Rank #86 Confirmed

Kimi K2.5 Instant Moonshot AI

43.6

Rank #89 Confirmed

Summary

  • They share 18 benchmarks with published results for both. Grok 4.3 scores higher in 3 categories and Kimi K2.5 Instant in 6 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Grok 4.3 leads 52.5 to 40.2.
  • Kimi K2.5 Instant has downloadable open weights; the other is API-only.

Side by side

Grok 4.3 and Kimi K2.5 Instant specifications
Grok 4.3Kimi K2.5 Instant
ProviderxAIMoonshot AI
Noometry Index43.843.6
Released2026-04-17—
WeightsProprietaryOpen
Context window1M—
Max output30K—
Input $ / M tokens$1.25—
Output $ / M tokens$2.50—
Results tracked4018

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

Grok 4.3: 41.6 (#121), Kimi K2.5 Instant: 42.6 (#97)

Coding benchmarks
BenchmarkGrok 4.3Kimi K2.5 Instant
LMArena WebDev13571404
LMArena Coding14151484
SciCode47.3%—
WeirdML49.9%—
ALE-Bench944.17—

Agentic & Tool Use Not comparable

Grok 4.3: 27.7 (#99), Kimi K2.5 Instant: —

Agentic & Tool Use benchmarks
BenchmarkGrok 4.3Kimi K2.5 Instant
GDP.pdf8%—
LMArena Search1165—
Vending-Bench 235.26—

Reasoning Grok 4.3 leads

Grok 4.3: 35.9 (#68), Kimi K2.5 Instant: 29.7 (#90)

Reasoning benchmarks
BenchmarkGrok 4.3Kimi K2.5 Instant
LMArena Hard Prompts13961443
NYT Connections (extended)55.2%—
CritPt8%—
Chess Puzzles25%—
DTBench90.7%—
LMCA38.3%—
Epoch Capabilities Index149.16—
ForecastBench60.3—

Math Grok 4.3 leads

Grok 4.3: 46.0 (#74), Kimi K2.5 Instant: 39.4 (#105)

Math benchmarks
BenchmarkGrok 4.3Kimi K2.5 Instant
LMArena Math13881442
FrontierMath (Tiers 1-3)42.8%—
FrontierMath Tier 414.6%—
OTIS Mock AIME 2024-202593.3%—
ProofBench11%—

Knowledge Grok 4.3 leads

Grok 4.3: 52.5 (#62), Kimi K2.5 Instant: 40.2 (#123)

Knowledge benchmarks
BenchmarkGrok 4.3Kimi K2.5 Instant
LMArena Expert13851440
GPQA Diamond88.8%—
SimpleQA Verified33.2%—

Multimodal Kimi K2.5 Instant leads

Grok 4.3: 31.6 (#104), Kimi K2.5 Instant: 40.2 (#50)

Multimodal benchmarks
BenchmarkGrok 4.3Kimi K2.5 Instant
LMArena Vision12291254
Blueprint-Bench 20%—

Multilingual Kimi K2.5 Instant leads

Grok 4.3: 50.5 (#120), Kimi K2.5 Instant: 52.0 (#94)

Multilingual benchmarks
BenchmarkGrok 4.3Kimi K2.5 Instant
LMArena Non-English13851406
LMArena Chinese14221449
LMArena French14121403
LMArena German13951413
LMArena Korean13561378
LMArena Russian13991404
LMArena Spanish13981447
LMArena Japanese1379—

Instruction Following Kimi K2.5 Instant leads

Grok 4.3: 72.1 (#140), Kimi K2.5 Instant: 75.3 (#65)

Instruction Following benchmarks
BenchmarkGrok 4.3Kimi K2.5 Instant
LMArena Instruction Following13661430

Long Context Kimi K2.5 Instant leads

Grok 4.3: 42.5 (#123), Kimi K2.5 Instant: 43.9 (#83)

Long Context benchmarks
BenchmarkGrok 4.3Kimi K2.5 Instant
LMArena Longer Query13931435

Writing & Preference Kimi K2.5 Instant leads

Grok 4.3: 58.5 (#118), Kimi K2.5 Instant: 60.6 (#95)

Writing & Preference benchmarks
BenchmarkGrok 4.3Kimi K2.5 Instant
LMArena Text13971420
LMArena Creative Writing13801381
LMArena Multi-Turn14061427
EQ-Bench 41075—

Frequently asked questions

Is Grok 4.3 better than Kimi K2.5 Instant?

Grok 4.3 and Kimi K2.5 Instant score almost the same on the Noometry Index (43.8 vs 43.6), so choose on price, context window or the category you care about most.

Is Grok 4.3 or Kimi K2.5 Instant better for coding?

They score almost the same on coding (41.6 vs 42.6); test both on your own repository before choosing.

How many benchmarks do Grok 4.3 and Kimi K2.5 Instant share?

18 benchmarks have published results for both models. Grok 4.3 has 40 scored results on Noometry and Kimi K2.5 Instant has 18.

Related comparisons

Go deeper