Model comparison

Kimi K2.5 Instant vs Qwen3 Max

Kimi K2.5 Instant and Qwen3 Max score almost the same on the Noometry Index (43.6 vs 43.7), so choose on price, context window or the category you care about most.

Last verified . 16 shared benchmarks.

Kimi K2.5 Instant Moonshot AI

43.6

Rank #89 Confirmed

Qwen3 Max Alibaba (Qwen)

43.7

Rank #87 Confirmed

Summary

  • They share 16 benchmarks with published results for both. Kimi K2.5 Instant scores higher in 4 categories and Qwen3 Max in 4 categories; 5 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Qwen3 Max leads 48.1 to 40.2.
  • Kimi K2.5 Instant has downloadable open weights; the other is API-only.

Side by side

Kimi K2.5 Instant and Qwen3 Max specifications
Kimi K2.5 InstantQwen3 Max
ProviderMoonshot AIAlibaba (Qwen)
Noometry Index43.643.7
Released—2025-09-23
WeightsOpenProprietary
Context window—262K
Max output—66K
Input $ / M tokens—$1.20
Output $ / M tokens—$6
Results tracked1833

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

Kimi K2.5 Instant: 42.6 (#97), Qwen3 Max: 43.0 (#93)

Coding benchmarks
BenchmarkKimi K2.5 InstantQwen3 Max
LMArena Coding14841456
LMArena WebDev1404—
ALE-Bench—370.45

Agentic & Tool Use Not comparable

Kimi K2.5 Instant: —, Qwen3 Max: —

Agentic & Tool Use benchmarks
BenchmarkKimi K2.5 InstantQwen3 Max
Vending-Bench 2—71.56

Reasoning Kimi K2.5 Instant leads

Kimi K2.5 Instant: 29.7 (#90), Qwen3 Max: 22.6 (#190)

Reasoning benchmarks
BenchmarkKimi K2.5 InstantQwen3 Max
LMArena Hard Prompts14431448
Kagi LLM Benchmark—72.5%
NYT Connections (extended)—30.1%
Chess Puzzles—4%
Mystery Game Puzzles—5%
DTBench—82.1%
LMCA—28.3%
Epoch Capabilities Index—142.38

Math Too close to call

Kimi K2.5 Instant: 39.4 (#105), Qwen3 Max: 38.7 (#131)

Math benchmarks
BenchmarkKimi K2.5 InstantQwen3 Max
LMArena Math14421446
FrontierMath (Tiers 1-3)—18.9%
OTIS Mock AIME 2024-2025—73.3%
MATH Level 5—97.1%

Knowledge Qwen3 Max leads

Kimi K2.5 Instant: 40.2 (#123), Qwen3 Max: 48.1 (#78)

Knowledge benchmarks
BenchmarkKimi K2.5 InstantQwen3 Max
LMArena Expert14401455
GPQA Diamond—72.6%
SimpleQA Verified—48.7%

Multimodal Not comparable

Kimi K2.5 Instant: 40.2 (#50), Qwen3 Max: —

Multimodal benchmarks
BenchmarkKimi K2.5 InstantQwen3 Max
LMArena Vision1254—

Multilingual Qwen3 Max leads

Kimi K2.5 Instant: 52.0 (#94), Qwen3 Max: 53.7 (#62)

Multilingual benchmarks
BenchmarkKimi K2.5 InstantQwen3 Max
LMArena Non-English14061429
LMArena Chinese14491478
LMArena French14031449
LMArena German14131463
LMArena Korean13781399
LMArena Russian14041428
LMArena Spanish14471462
LMArena Japanese—1397

Instruction Following Too close to call

Kimi K2.5 Instant: 75.3 (#65), Qwen3 Max: 74.8 (#87)

Instruction Following benchmarks
BenchmarkKimi K2.5 InstantQwen3 Max
LMArena Instruction Following14301419

Long Context Kimi K2.5 Instant leads

Kimi K2.5 Instant: 43.9 (#83), Qwen3 Max: 41.6 (#134)

Long Context benchmarks
BenchmarkKimi K2.5 InstantQwen3 Max
LMArena Longer Query14351438
Fiction.LiveBench—66.7%
CL-bench—14.5%

Writing & Preference Qwen3 Max leads

Kimi K2.5 Instant: 60.6 (#95), Qwen3 Max: 62.4 (#76)

Writing & Preference benchmarks
BenchmarkKimi K2.5 InstantQwen3 Max
LMArena Text14201439
LMArena Creative Writing13811402
LMArena Multi-Turn14271446

Frequently asked questions

Is Kimi K2.5 Instant better than Qwen3 Max?

Kimi K2.5 Instant and Qwen3 Max score almost the same on the Noometry Index (43.6 vs 43.7), so choose on price, context window or the category you care about most.

Is Kimi K2.5 Instant or Qwen3 Max better for coding?

They score almost the same on coding (42.6 vs 43.0); test both on your own repository before choosing.

How many benchmarks do Kimi K2.5 Instant and Qwen3 Max share?

16 benchmarks have published results for both models. Kimi K2.5 Instant has 18 scored results on Noometry and Qwen3 Max has 33.

Related comparisons

Go deeper