Model comparison

Claude Opus 4.8 vs Kimi K3

Claude Opus 4.8 is the stronger model overall, scoring 60.7 to 59.5 on the Noometry Index. Kimi K3 costs 1.7× less per token, which makes it the better buy when Claude Opus 4.8's lead doesn't matter for your workload.

Last verified . 52 shared benchmarks.

Claude Opus 4.8 Anthropic

60.7

Rank #13 Confirmed

Kimi K3 Moonshot AI

59.5

Rank #15 Confirmed

Summary

  • They share 52 benchmarks with published results for both. Claude Opus 4.8 scores higher in 4 categories and Kimi K3 in 6 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in agentic & tool use, where Claude Opus 4.8 leads 47.6 to 41.8.
  • The biggest single-benchmark swing is GBAEval: 70.9% for Claude Opus 4.8 and 48.3% for Kimi K3.
  • Kimi K3 is cheaper at $3 / $15 per million input/output tokens, against $5 / $25 for Claude Opus 4.8.
  • Kimi K3 accepts more context: 1.05M tokens versus 1M.
  • Kimi K3 has downloadable open weights; the other is API-only.

Side by side

Claude Opus 4.8 and Kimi K3 specifications
Claude Opus 4.8Kimi K3
ProviderAnthropicMoonshot AI
Noometry Index60.759.5
Released2026-05-282026-07-16
WeightsProprietaryOpen
Context window1M1.05M
Max output128K1.05M
Input $ / M tokens$5$3
Output $ / M tokens$25$15
Results tracked6553

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Kimi K3 leads

Claude Opus 4.8: 59.9 (#12), Kimi K3: 61.0 (#10)

Coding benchmarks
BenchmarkClaude Opus 4.8Kimi K3
DeepSWE59%68.5%
FrontierCode46.5%44.2%
LMArena WebDev15561654
SciCode53.5%59.5%
WeirdML82.9%82.6%
LMArena Coding14901508
ALE-Bench1,5641,524
FrontierSWE—25.9%
GSO47.1%—

Agentic & Tool Use Claude Opus 4.8 leads

Claude Opus 4.8: 47.6 (#11), Kimi K3: 41.8 (#20)

Agentic & Tool Use benchmarks
BenchmarkClaude Opus 4.8Kimi K3
APEX-Agents48.9%50.6%
τ²-bench Banking39.7%37.1%
PostTrainBench33.8%32%
GBAEval70.9%48.3%
GDP.pdf24%19%
Vending-Bench 25,7875,165
OSWorld 2.020.6%—
Remote Labor Index8.3%—
DeepResearch Bench50.2%—
LMArena Search1204—

Reasoning Claude Opus 4.8 leads

Claude Opus 4.8: 64.7 (#16), Kimi K3: 63.0 (#17)

Reasoning benchmarks
BenchmarkClaude Opus 4.8Kimi K3
ARC-AGI-272.1%60.4%
SimpleBench64.8%60.7%
NYT Connections (extended)91.1%93.6%
ARC-AGI-192.5%94.5%
CritPt20.9%23.4%
Chess Puzzles34%39%
LMArena Hard Prompts14821496
Mystery Game Puzzles36%26%
DTBench94.9%91.2%
LMCA57.5%52.7%
Surface Evolver Bench87.5%95%
Epoch Capabilities Index158.21157.45
ForecastBench59.961.1
Kagi LLM Benchmark88.8%—
EnigmaEval23.5%—
EBR-Bench28.6%—
Bench to the Future 30.14—

Math Claude Opus 4.8 leads

Claude Opus 4.8: 78.4 (#13), Kimi K3: 74.2 (#16)

Knowledge Kimi K3 leads

Claude Opus 4.8: 61.3 (#29), Kimi K3: 63.2 (#21)

Knowledge benchmarks
BenchmarkClaude Opus 4.8Kimi K3
GPQA Diamond91%93.1%
SimpleQA Verified53%50.6%
LMArena Expert15021521

Multimodal Claude Opus 4.8 leads

Claude Opus 4.8: 42.9 (#26), Kimi K3: 37.8 (#70)

Multimodal benchmarks
BenchmarkClaude Opus 4.8Kimi K3
Blueprint-Bench 214.5%29.5%
Furniture Assembly42.5%34.2%
LMArena Vision1294—
LMArena Document1475—

Multilingual Kimi K3 leads

Claude Opus 4.8: 55.2 (#33), Kimi K3: 56.3 (#21)

Multilingual benchmarks
BenchmarkClaude Opus 4.8Kimi K3
LMArena Non-English14501466
LMArena Chinese15071529
LMArena French14811491
LMArena German14721488
LMArena Japanese14401487
LMArena Korean14321458
LMArena Russian14741482
LMArena Spanish14661472

Instruction Following Too close to call

Claude Opus 4.8: 77.4 (#24), Kimi K3: 77.7 (#14)

Instruction Following benchmarks
BenchmarkClaude Opus 4.8Kimi K3
LMArena Instruction Following14761483

Long Context Too close to call

Claude Opus 4.8: 45.4 (#35), Kimi K3: 45.8 (#29)

Long Context benchmarks
BenchmarkClaude Opus 4.8Kimi K3
LMArena Longer Query14831494

Writing & Preference Kimi K3 leads

Claude Opus 4.8: 72.0 (#16), Kimi K3: 76.6 (#4)

Writing & Preference benchmarks
BenchmarkClaude Opus 4.8Kimi K3
LMArena Text14611476
LMArena Creative Writing14541454
EQ-Bench Creative Writing18402082
EQ-Bench 412811339
LMArena Multi-Turn14761488

Frequently asked questions

Is Claude Opus 4.8 better than Kimi K3?

Claude Opus 4.8 is the stronger model overall, scoring 60.7 to 59.5 on the Noometry Index. Kimi K3 costs 1.7× less per token, which makes it the better buy when Claude Opus 4.8's lead doesn't matter for your workload.

Which is cheaper, Claude Opus 4.8 or Kimi K3?

Kimi K3 is cheaper. It lists at $3 per million input tokens and $15 per million output tokens; Claude Opus 4.8 lists at $5 and $25.

Is Claude Opus 4.8 or Kimi K3 better for coding?

Kimi K3 scores higher on coding benchmarks: 61.0 versus 59.9 in the Noometry coding category.

Which has the bigger context window?

Kimi K3 does, with 1.05M tokens against 1M.

How many benchmarks do Claude Opus 4.8 and Kimi K3 share?

52 benchmarks have published results for both models. Claude Opus 4.8 has 65 scored results on Noometry and Kimi K3 has 53.

Related comparisons

Go deeper