Model comparison

Kimi K3 vs Qwen3.5 Max Preview

Kimi K3 is the stronger model overall, scoring 59.5 to 45.3 on the Noometry Index.

Last verified . 17 shared benchmarks.

Kimi K3 Moonshot AI

59.5

Rank #15 Confirmed

Qwen3.5 Max Preview Alibaba (Qwen)

45.3

Rank #71 Confirmed

Summary

  • They share 17 benchmarks with published results for both. Kimi K3 scores higher in 8 categories and Qwen3.5 Max Preview in 0 categories; 5 gaps are clear of the uncertainty.
  • The widest gap is in math, where Kimi K3 leads 74.2 to 40.1.
  • Kimi K3 has downloadable open weights; the other is API-only.

Side by side

Kimi K3 and Qwen3.5 Max Preview specifications
Kimi K3Qwen3.5 Max Preview
ProviderMoonshot AIAlibaba (Qwen)
Noometry Index59.545.3
Released2026-07-16—
WeightsOpenProprietary
Context window1.05M—
Max output1.05M—
Input $ / M tokens$3—
Output $ / M tokens$15—
Results tracked5317

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Kimi K3 leads

Kimi K3: 61.0 (#10), Qwen3.5 Max Preview: 44.0 (#77)

Coding benchmarks
BenchmarkKimi K3Qwen3.5 Max Preview
LMArena Coding15081487
DeepSWE68.5%—
FrontierCode44.2%—
LMArena WebDev1654—
FrontierSWE25.9%—
SciCode59.5%—
WeirdML82.6%—
ALE-Bench1,524—

Agentic & Tool Use Not comparable

Kimi K3: 41.8 (#20), Qwen3.5 Max Preview: —

Agentic & Tool Use benchmarks
BenchmarkKimi K3Qwen3.5 Max Preview
APEX-Agents50.6%—
τ²-bench Banking37.1%—
PostTrainBench32%—
GBAEval48.3%—
GDP.pdf19%—
Vending-Bench 25,165—

Reasoning Kimi K3 leads

Kimi K3: 63.0 (#17), Qwen3.5 Max Preview: 30.8 (#84)

Reasoning benchmarks
BenchmarkKimi K3Qwen3.5 Max Preview
LMArena Hard Prompts14961483
ARC-AGI-260.4%—
SimpleBench60.7%—
NYT Connections (extended)93.6%—
ARC-AGI-194.5%—
CritPt23.4%—
Chess Puzzles39%—
Mystery Game Puzzles26%—
DTBench91.2%—
LMCA52.7%—
Surface Evolver Bench95%—
Epoch Capabilities Index157.45—
ForecastBench61.1—

Math Kimi K3 leads

Kimi K3: 74.2 (#16), Qwen3.5 Max Preview: 40.1 (#94)

Math benchmarks
BenchmarkKimi K3Qwen3.5 Max Preview
LMArena Math14911474
FrontierMath (Tiers 1-3)72.2%—
FrontierMath Tier 439%—
MathArena Final-Answer Competitions87.8%—
OTIS Mock AIME 2024-202597.2%—
ProofBench87%—

Knowledge Kimi K3 leads

Kimi K3: 63.2 (#21), Qwen3.5 Max Preview: 41.8 (#107)

Knowledge benchmarks
BenchmarkKimi K3Qwen3.5 Max Preview
LMArena Expert15211489
GPQA Diamond93.1%—
SimpleQA Verified50.6%—

Multimodal Not comparable

Kimi K3: 37.8 (#70), Qwen3.5 Max Preview: —

Multimodal benchmarks
BenchmarkKimi K3Qwen3.5 Max Preview
Blueprint-Bench 229.5%—
Furniture Assembly34.2%—

Multilingual Too close to call

Kimi K3: 56.3 (#21), Qwen3.5 Max Preview: 56.2 (#22)

Multilingual benchmarks
BenchmarkKimi K3Qwen3.5 Max Preview
LMArena Non-English14661465
LMArena Chinese15291534
LMArena French14911484
LMArena German14881487
LMArena Japanese14871495
LMArena Korean14581438
LMArena Russian14821471
LMArena Spanish14721470

Instruction Following Too close to call

Kimi K3: 77.7 (#14), Qwen3.5 Max Preview: 77.0 (#31)

Instruction Following benchmarks
BenchmarkKimi K3Qwen3.5 Max Preview
LMArena Instruction Following14831467

Long Context Too close to call

Kimi K3: 45.8 (#29), Qwen3.5 Max Preview: 45.2 (#45)

Long Context benchmarks
BenchmarkKimi K3Qwen3.5 Max Preview
LMArena Longer Query14941476

Writing & Preference Kimi K3 leads

Kimi K3: 76.6 (#4), Qwen3.5 Max Preview: 66.0 (#41)

Writing & Preference benchmarks
BenchmarkKimi K3Qwen3.5 Max Preview
LMArena Text14761470
LMArena Creative Writing14541464
LMArena Multi-Turn14881478
EQ-Bench Creative Writing2082—
EQ-Bench 41339—

Frequently asked questions

Is Kimi K3 better than Qwen3.5 Max Preview?

Kimi K3 is the stronger model overall, scoring 59.5 to 45.3 on the Noometry Index.

Is Kimi K3 or Qwen3.5 Max Preview better for coding?

Kimi K3 scores higher on coding benchmarks: 61.0 versus 44.0 in the Noometry coding category.

How many benchmarks do Kimi K3 and Qwen3.5 Max Preview share?

17 benchmarks have published results for both models. Kimi K3 has 53 scored results on Noometry and Qwen3.5 Max Preview has 17.

Related comparisons

Go deeper