Model comparison

Kimi K3 vs Qwen3 Max

Kimi K3 is the stronger model overall, scoring 59.5 to 43.7 on the Noometry Index. Qwen3 Max costs 2.5× less per token, which makes it the better buy when Kimi K3's lead doesn't matter for your workload.

Last verified . 29 shared benchmarks.

Kimi K3 Moonshot AI

59.5

Rank #15 Confirmed

Qwen3 Max Alibaba (Qwen)

43.7

Rank #87 Confirmed

Summary

  • They share 29 benchmarks with published results for both. Kimi K3 scores higher in 8 categories and Qwen3 Max in 0 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Kimi K3 leads 63.0 to 22.6.
  • The biggest single-benchmark swing is NYT Connections (extended): 93.6% for Kimi K3 and 30.1% for Qwen3 Max.
  • Qwen3 Max is cheaper at $1.20 / $6 per million input/output tokens, against $3 / $15 for Kimi K3.
  • Kimi K3 accepts more context: 1.05M tokens versus 262K.
  • Kimi K3 has downloadable open weights; the other is API-only.

Side by side

Kimi K3 and Qwen3 Max specifications
Kimi K3Qwen3 Max
ProviderMoonshot AIAlibaba (Qwen)
Noometry Index59.543.7
Released2026-07-162025-09-23
WeightsOpenProprietary
Context window1.05M262K
Max output1.05M66K
Input $ / M tokens$3$1.20
Output $ / M tokens$15$6
Results tracked5333

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Kimi K3 leads

Kimi K3: 61.0 (#10), Qwen3 Max: 43.0 (#93)

Coding benchmarks
BenchmarkKimi K3Qwen3 Max
LMArena Coding15081456
ALE-Bench1,524370.45
DeepSWE68.5%—
FrontierCode44.2%—
LMArena WebDev1654—
FrontierSWE25.9%—
SciCode59.5%—
WeirdML82.6%—

Agentic & Tool Use Not comparable

Kimi K3: 41.8 (#20), Qwen3 Max: —

Agentic & Tool Use benchmarks
BenchmarkKimi K3Qwen3 Max
Vending-Bench 25,16571.56
APEX-Agents50.6%—
τ²-bench Banking37.1%—
PostTrainBench32%—
GBAEval48.3%—
GDP.pdf19%—

Reasoning Kimi K3 leads

Kimi K3: 63.0 (#17), Qwen3 Max: 22.6 (#190)

Reasoning benchmarks
BenchmarkKimi K3Qwen3 Max
NYT Connections (extended)93.6%30.1%
Chess Puzzles39%4%
LMArena Hard Prompts14961448
Mystery Game Puzzles26%5%
DTBench91.2%82.1%
LMCA52.7%28.3%
Epoch Capabilities Index157.45142.38
ARC-AGI-260.4%—
SimpleBench60.7%—
Kagi LLM Benchmark—72.5%
ARC-AGI-194.5%—
CritPt23.4%—
Surface Evolver Bench95%—
ForecastBench61.1—

Math Kimi K3 leads

Kimi K3: 74.2 (#16), Qwen3 Max: 38.7 (#131)

Math benchmarks
BenchmarkKimi K3Qwen3 Max
FrontierMath (Tiers 1-3)72.2%18.9%
OTIS Mock AIME 2024-202597.2%73.3%
LMArena Math14911446
FrontierMath Tier 439%—
MathArena Final-Answer Competitions87.8%—
ProofBench87%—
MATH Level 5—97.1%

Knowledge Kimi K3 leads

Kimi K3: 63.2 (#21), Qwen3 Max: 48.1 (#78)

Knowledge benchmarks
BenchmarkKimi K3Qwen3 Max
GPQA Diamond93.1%72.6%
SimpleQA Verified50.6%48.7%
LMArena Expert15211455

Multimodal Not comparable

Kimi K3: 37.8 (#70), Qwen3 Max: —

Multimodal benchmarks
BenchmarkKimi K3Qwen3 Max
Blueprint-Bench 229.5%—
Furniture Assembly34.2%—

Multilingual Kimi K3 leads

Kimi K3: 56.3 (#21), Qwen3 Max: 53.7 (#62)

Multilingual benchmarks
BenchmarkKimi K3Qwen3 Max
LMArena Non-English14661429
LMArena Chinese15291478
LMArena French14911449
LMArena German14881463
LMArena Japanese14871397
LMArena Korean14581399
LMArena Russian14821428
LMArena Spanish14721462

Instruction Following Kimi K3 leads

Kimi K3: 77.7 (#14), Qwen3 Max: 74.8 (#87)

Instruction Following benchmarks
BenchmarkKimi K3Qwen3 Max
LMArena Instruction Following14831419

Long Context Kimi K3 leads

Kimi K3: 45.8 (#29), Qwen3 Max: 41.6 (#134)

Long Context benchmarks
BenchmarkKimi K3Qwen3 Max
LMArena Longer Query14941438
Fiction.LiveBench—66.7%
CL-bench—14.5%

Writing & Preference Kimi K3 leads

Kimi K3: 76.6 (#4), Qwen3 Max: 62.4 (#76)

Writing & Preference benchmarks
BenchmarkKimi K3Qwen3 Max
LMArena Text14761439
LMArena Creative Writing14541402
LMArena Multi-Turn14881446
EQ-Bench Creative Writing2082—
EQ-Bench 41339—

Frequently asked questions

Is Kimi K3 better than Qwen3 Max?

Kimi K3 is the stronger model overall, scoring 59.5 to 43.7 on the Noometry Index. Qwen3 Max costs 2.5× less per token, which makes it the better buy when Kimi K3's lead doesn't matter for your workload.

Which is cheaper, Kimi K3 or Qwen3 Max?

Qwen3 Max is cheaper. It lists at $1.20 per million input tokens and $6 per million output tokens; Kimi K3 lists at $3 and $15.

Is Kimi K3 or Qwen3 Max better for coding?

Kimi K3 scores higher on coding benchmarks: 61.0 versus 43.0 in the Noometry coding category.

Which has the bigger context window?

Kimi K3 does, with 1.05M tokens against 262K.

How many benchmarks do Kimi K3 and Qwen3 Max share?

29 benchmarks have published results for both models. Kimi K3 has 53 scored results on Noometry and Qwen3 Max has 33.

Related comparisons

Go deeper