Model comparison

GPT-5.4 vs Kimi K3

GPT-5.4 and Kimi K3 score almost the same on the Noometry Index (59.4 vs 59.5), so choose on price, context window or the category you care about most.

Last verified . 48 shared benchmarks.

GPT-5.4 OpenAI

59.4

Rank #16 Confirmed

Kimi K3 Moonshot AI

59.5

Rank #15 Confirmed

Summary

  • They share 48 benchmarks with published results for both. GPT-5.4 scores higher in 4 categories and Kimi K3 in 6 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in coding, where Kimi K3 leads 61.0 to 52.6.
  • The biggest single-benchmark swing is ProofBench: 56% for GPT-5.4 and 87% for Kimi K3.
  • GPT-5.4 is cheaper at $2.50 / $15 per million input/output tokens, against $3 / $15 for Kimi K3.
  • GPT-5.4 accepts more context: 1.05M tokens versus 1.05M.
  • Kimi K3 has downloadable open weights; the other is API-only.

Side by side

GPT-5.4 and Kimi K3 specifications
GPT-5.4Kimi K3
ProviderOpenAIMoonshot AI
Noometry Index59.459.5
Released2026-03-052026-07-16
WeightsProprietaryOpen
Context window1.05M1.05M
Max output128K1.05M
Input $ / M tokens$2.50$3
Output $ / M tokens$15$15
Results tracked6853

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Kimi K3 leads

GPT-5.4: 52.6 (#33), Kimi K3: 61.0 (#10)

Coding benchmarks
BenchmarkGPT-5.4Kimi K3
DeepSWE51.8%68.5%
LMArena WebDev14651654
SciCode56.6%59.5%
WeirdML77.7%82.6%
LMArena Coding14971508
ALE-Bench1,6071,524
SWE-bench Verified76.9%—
FrontierCode—44.2%
FrontierSWE—25.9%
GSO31.4%—
MirrorCode15.6%—
AlgoTune1.85—

Agentic & Tool Use GPT-5.4 leads

GPT-5.4: 46.5 (#13), Kimi K3: 41.8 (#20)

Agentic & Tool Use benchmarks
BenchmarkGPT-5.4Kimi K3
APEX-Agents52.4%50.6%
τ²-bench Banking39.4%37.1%
PostTrainBench19%32%
GBAEval45.1%48.3%
Vending-Bench 26,1445,165
Terminal-Bench81.8%—
DeepResearch Bench35.1%—
GDP.pdf—19%
LMArena Search1197—
METR Time Horizons74.3%—

Reasoning Kimi K3 leads

GPT-5.4: 61.8 (#19), Kimi K3: 63.0 (#17)

Reasoning benchmarks
BenchmarkGPT-5.4Kimi K3
ARC-AGI-274%60.4%
NYT Connections (extended)91.3%93.6%
ARC-AGI-193.7%94.5%
CritPt23.4%23.4%
Chess Puzzles44%39%
LMArena Hard Prompts14851496
Mystery Game Puzzles37%26%
DTBench94.4%91.2%
LMCA52%52.7%
Epoch Capabilities Index156.81157.45
ForecastBench59.561.1
SimpleBench—60.7%
Kagi LLM Benchmark63.8%—
EnigmaEval16%—
Thematic Generalization80%—
EBR-Bench25.4%—
Surface Evolver Bench—95%

Math Too close to call

GPT-5.4: 73.5 (#19), Kimi K3: 74.2 (#16)

Knowledge GPT-5.4 leads

GPT-5.4: 65.3 (#14), Kimi K3: 63.2 (#21)

Knowledge benchmarks
BenchmarkGPT-5.4Kimi K3
GPQA Diamond93.3%93.1%
SimpleQA Verified45.1%50.6%
LMArena Expert15071521
Humanity's Last Exam36.2%—
Vectara Hallucination Rate7%—

Multimodal GPT-5.4 leads

GPT-5.4: 43.7 (#20), Kimi K3: 37.8 (#70)

Multimodal benchmarks
BenchmarkGPT-5.4Kimi K3
Blueprint-Bench 227.1%29.5%
Furniture Assembly37.5%34.2%
LMArena Vision1303—
LMArena Document1471—

Multilingual Too close to call

GPT-5.4: 56.2 (#23), Kimi K3: 56.3 (#21)

Multilingual benchmarks
BenchmarkGPT-5.4Kimi K3
LMArena Non-English14651466
LMArena Chinese15191529
LMArena French14931491
LMArena German14721488
LMArena Japanese14851487
LMArena Korean14481458
LMArena Russian14801482
LMArena Spanish14541472

Instruction Following Too close to call

GPT-5.4: 77.1 (#27), Kimi K3: 77.7 (#14)

Instruction Following benchmarks
BenchmarkGPT-5.4Kimi K3
LMArena Instruction Following14691483

Long Context GPT-5.4 leads

GPT-5.4: 50.3 (#8), Kimi K3: 45.8 (#29)

Long Context benchmarks
BenchmarkGPT-5.4Kimi K3
LMArena Longer Query14731494
CL-bench27.9%—
CL-bench Life21.7%—

Writing & Preference Kimi K3 leads

GPT-5.4: 71.9 (#17), Kimi K3: 76.6 (#4)

Writing & Preference benchmarks
BenchmarkGPT-5.4Kimi K3
LMArena Text14691476
LMArena Creative Writing14391454
EQ-Bench Creative Writing18402082
EQ-Bench 412721339
LMArena Multi-Turn14821488

Frequently asked questions

Is GPT-5.4 better than Kimi K3?

GPT-5.4 and Kimi K3 score almost the same on the Noometry Index (59.4 vs 59.5), so choose on price, context window or the category you care about most.

Which is cheaper, GPT-5.4 or Kimi K3?

GPT-5.4 is cheaper. It lists at $2.50 per million input tokens and $15 per million output tokens; Kimi K3 lists at $3 and $15.

Is GPT-5.4 or Kimi K3 better for coding?

Kimi K3 scores higher on coding benchmarks: 61.0 versus 52.6 in the Noometry coding category.

Which has the bigger context window?

GPT-5.4 does, with 1.05M tokens against 1.05M.

How many benchmarks do GPT-5.4 and Kimi K3 share?

48 benchmarks have published results for both models. GPT-5.4 has 68 scored results on Noometry and Kimi K3 has 53.

Related comparisons

Go deeper