Model comparison

Claude Opus 4 vs Kimi K3

Kimi K3 is the stronger model overall, scoring 59.5 to 43.1 on the Noometry Index.

Last verified . 29 shared benchmarks.

Claude Opus 4 Anthropic

43.1

Rank #100 Confirmed

Kimi K3 Moonshot AI

59.5

Rank #15 Confirmed

Summary

  • They share 29 benchmarks with published results for both. Claude Opus 4 scores higher in 0 categories and Kimi K3 in 10 categories; 9 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Kimi K3 leads 63.0 to 27.3.
  • The biggest single-benchmark swing is ARC-AGI-1: 35.7% for Claude Opus 4 and 94.5% for Kimi K3.
  • Kimi K3 is cheaper at $3 / $15 per million input/output tokens, against $15 / $75 for Claude Opus 4.
  • Kimi K3 accepts more context: 1.05M tokens versus 200K.
  • Kimi K3 has downloadable open weights; the other is API-only.

Side by side

Claude Opus 4 and Kimi K3 specifications
Claude Opus 4Kimi K3
ProviderAnthropicMoonshot AI
Noometry Index43.159.5
Released2025-05-222026-07-16
WeightsProprietaryOpen
Context window200K1.05M
Max output32K1.05M
Input $ / M tokens$15$3
Output $ / M tokens$75$15
Results tracked5653

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Kimi K3 leads

Claude Opus 4: 47.2 (#62), Kimi K3: 61.0 (#10)

Coding benchmarks
BenchmarkClaude Opus 4Kimi K3
WeirdML43.7%82.6%
LMArena Coding14421508
SWE-bench Verified70.7%—
DeepSWE—68.5%
FrontierCode—44.2%
SWE-bench Verified (bash only)67.6%—
Aider Polyglot72%—
LMArena WebDev—1654
FrontierSWE—25.9%
SciCode—59.5%
GSO6.9%—
ALE-Bench—1,524
AlgoTune1.33—

Agentic & Tool Use Kimi K3 leads

Claude Opus 4: 34.8 (#42), Kimi K3: 41.8 (#20)

Agentic & Tool Use benchmarks
BenchmarkClaude Opus 4Kimi K3
APEX-Agents—50.6%
τ²-bench Banking—37.1%
Cybench38%—
DeepResearch Bench46.8%—
PostTrainBench—32%
GBAEval—48.3%
GDP.pdf—19%
LMArena Search1127—
METR Time Horizons63.9%—
Vending-Bench 2—5,165

Reasoning Kimi K3 leads

Claude Opus 4: 27.3 (#121), Kimi K3: 63.0 (#17)

Reasoning benchmarks
BenchmarkClaude Opus 4Kimi K3
ARC-AGI-28.6%60.4%
SimpleBench58.8%60.7%
ARC-AGI-135.7%94.5%
CritPt0.3%23.4%
LMArena Hard Prompts13991496
DTBench81.6%91.2%
LMCA37.4%52.7%
Epoch Capabilities Index142.67157.45
ForecastBench61.161.1
Kagi LLM Benchmark74.3%—
NYT Connections (extended)—93.6%
Chess Puzzles—39%
EnigmaEval5.6%—
Mystery Game Puzzles—26%
Surface Evolver Bench—95%

Math Kimi K3 leads

Claude Opus 4: 42.0 (#86), Kimi K3: 74.2 (#16)

Knowledge Kimi K3 leads

Claude Opus 4: 44.0 (#88), Kimi K3: 63.2 (#21)

Knowledge benchmarks
BenchmarkClaude Opus 4Kimi K3
GPQA Diamond76.3%93.1%
LMArena Expert13861521
Humanity's Last Exam10.7%—
SimpleQA Verified—50.6%
MMLU-Pro87.5%—
Confabulations15.9%—
Vectara Hallucination Rate12%—
GPQA (HELM)70.8%—

Multimodal Kimi K3 leads

Claude Opus 4: 31.5 (#106), Kimi K3: 37.8 (#70)

Multimodal benchmarks
BenchmarkClaude Opus 4Kimi K3
LMArena Vision1192—
GeoBench49%—
VPCT38%—
Blueprint-Bench 2—29.5%
Furniture Assembly—34.2%

Multilingual Kimi K3 leads

Claude Opus 4: 48.8 (#138), Kimi K3: 56.3 (#21)

Multilingual benchmarks
BenchmarkClaude Opus 4Kimi K3
LMArena Non-English13621466
LMArena Chinese13861529
LMArena French13721491
LMArena German13911488
LMArena Japanese13311487
LMArena Korean13211458
LMArena Russian13921482
LMArena Spanish13891472

Instruction Following Too close to call

Claude Opus 4: 77.1 (#28), Kimi K3: 77.7 (#14)

Instruction Following benchmarks
BenchmarkClaude Opus 4Kimi K3
LMArena Instruction Following14061483
IFEval91.8%—

Long Context Kimi K3 leads

Claude Opus 4: 39.6 (#172), Kimi K3: 45.8 (#29)

Long Context benchmarks
BenchmarkClaude Opus 4Kimi K3
LMArena Longer Query14221494
Fiction.LiveBench61.1%—

Writing & Preference Kimi K3 leads

Claude Opus 4: 61.2 (#89), Kimi K3: 76.6 (#4)

Writing & Preference benchmarks
BenchmarkClaude Opus 4Kimi K3
LMArena Text13771476
LMArena Creative Writing13871454
EQ-Bench Creative Writing15802082
LMArena Multi-Turn13961488
Short-Story Creative Writing83.6%—
WildBench85.2%—
EQ-Bench 4—1339

Frequently asked questions

Is Claude Opus 4 better than Kimi K3?

Kimi K3 is the stronger model overall, scoring 59.5 to 43.1 on the Noometry Index.

Which is cheaper, Claude Opus 4 or Kimi K3?

Kimi K3 is cheaper. It lists at $3 per million input tokens and $15 per million output tokens; Claude Opus 4 lists at $15 and $75.

Is Claude Opus 4 or Kimi K3 better for coding?

Kimi K3 scores higher on coding benchmarks: 61.0 versus 47.2 in the Noometry coding category.

Which has the bigger context window?

Kimi K3 does, with 1.05M tokens against 200K.

How many benchmarks do Claude Opus 4 and Kimi K3 share?

29 benchmarks have published results for both models. Claude Opus 4 has 56 scored results on Noometry and Kimi K3 has 53.

Related comparisons

Go deeper