Model comparison

Kimi K3 vs Qwen2.5 72B Instruct

Kimi K3 is the stronger model overall, scoring 59.5 to 31.9 on the Noometry Index. Qwen2.5 72B Instruct costs 2.4× less per token, which makes it the better buy when Kimi K3's lead doesn't matter for your workload.

Last verified . 24 shared benchmarks.

Kimi K3 Moonshot AI

59.5

Rank #15 Confirmed

Qwen2.5 72B Instruct Alibaba (Qwen)

31.9

Rank #267 Confirmed

Summary

  • They share 24 benchmarks with published results for both. Kimi K3 scores higher in 9 categories and Qwen2.5 72B Instruct in 0 categories; 9 gaps are clear of the uncertainty.
  • The widest gap is in math, where Kimi K3 leads 74.2 to 19.3.
  • The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 97.2% for Kimi K3 and 8.1% for Qwen2.5 72B Instruct.
  • Qwen2.5 72B Instruct is cheaper at $1.40 / $5.60 per million input/output tokens, against $3 / $15 for Kimi K3.
  • Kimi K3 accepts more context: 1.05M tokens versus 131K.

Side by side

Kimi K3 and Qwen2.5 72B Instruct specifications
Kimi K3Qwen2.5 72B Instruct
ProviderMoonshot AIAlibaba (Qwen)
Noometry Index59.531.9
Released2026-07-162024-09
WeightsOpenOpen
Context window1.05M131K
Max output1.05M8K
Input $ / M tokens$3$1.40
Output $ / M tokens$15$5.60
Results tracked5343

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Kimi K3 leads

Kimi K3: 61.0 (#10), Qwen2.5 72B Instruct: 33.2 (#260)

Coding benchmarks
BenchmarkKimi K3Qwen2.5 72B Instruct
WeirdML82.6%16%
LMArena Coding15081292
DeepSWE68.5%—
FrontierCode44.2%—
LMArena WebDev1654—
FrontierSWE25.9%—
SciCode59.5%—
BigCodeBench Instruct—45.8%
BigCodeBench Complete—55.9%
ALE-Bench1,524—

Agentic & Tool Use Kimi K3 leads

Kimi K3: 41.8 (#20), Qwen2.5 72B Instruct: 22.1 (#133)

Agentic & Tool Use benchmarks
BenchmarkKimi K3Qwen2.5 72B Instruct
APEX-Agents50.6%—
TheAgentCompany—5.7%
τ²-bench Banking37.1%—
PostTrainBench32%—
BALROG—16.2%
GBAEval48.3%—
GDP.pdf19%—
METR Time Horizons—35.8%
Vending-Bench 25,165—

Reasoning Kimi K3 leads

Kimi K3: 63.0 (#17), Qwen2.5 72B Instruct: 22.3 (#199)

Reasoning benchmarks
BenchmarkKimi K3Qwen2.5 72B Instruct
LMArena Hard Prompts14961271
DTBench91.2%62.9%
LMCA52.7%13.4%
Epoch Capabilities Index157.45129
ForecastBench61.157.5
ARC-AGI-260.4%—
SimpleBench60.7%—
NYT Connections (extended)93.6%—
ARC-AGI-194.5%—
CritPt23.4%—
Chess Puzzles39%—
Mystery Game Puzzles26%—
Surface Evolver Bench95%—
BIG-Bench Hard—79.8%
HellaSwag—84.8%
PIQA—82.6%
WinoGrande—82.3%

Math Kimi K3 leads

Kimi K3: 74.2 (#16), Qwen2.5 72B Instruct: 19.3 (#287)

Math benchmarks
BenchmarkKimi K3Qwen2.5 72B Instruct
OTIS Mock AIME 2024-202597.2%8.1%
LMArena Math14911283
FrontierMath (Tiers 1-3)72.2%—
FrontierMath Tier 439%—
MathArena Final-Answer Competitions87.8%—
ProofBench87%—
Omni-MATH—33%
MATH Level 5—63.2%

Knowledge Kimi K3 leads

Kimi K3: 63.2 (#21), Qwen2.5 72B Instruct: 27.0 (#253)

Knowledge benchmarks
BenchmarkKimi K3Qwen2.5 72B Instruct
GPQA Diamond93.1%49.1%
LMArena Expert15211245
SimpleQA Verified50.6%—
MMLU-Pro—63.1%
Confabulations—19.1%
GPQA (HELM)—42.6%
ARC (AI2) Challenge—94.5%
MMLU—85.3%
TriviaQA—71.9%

Multimodal Not comparable

Kimi K3: 37.8 (#70), Qwen2.5 72B Instruct: —

Multimodal benchmarks
BenchmarkKimi K3Qwen2.5 72B Instruct
Blueprint-Bench 229.5%—
Furniture Assembly34.2%—

Multilingual Kimi K3 leads

Kimi K3: 56.3 (#21), Qwen2.5 72B Instruct: 41.0 (#213)

Multilingual benchmarks
BenchmarkKimi K3Qwen2.5 72B Instruct
LMArena Non-English14661252
LMArena Chinese15291272
LMArena French14911280
LMArena German14881234
LMArena Japanese14871180
LMArena Korean14581188
LMArena Russian14821264
LMArena Spanish14721256

Instruction Following Kimi K3 leads

Kimi K3: 77.7 (#14), Qwen2.5 72B Instruct: 65.5 (#221)

Instruction Following benchmarks
BenchmarkKimi K3Qwen2.5 72B Instruct
LMArena Instruction Following14831254
IFEval—80.6%

Long Context Kimi K3 leads

Kimi K3: 45.8 (#29), Qwen2.5 72B Instruct: 38.9 (#188)

Long Context benchmarks
BenchmarkKimi K3Qwen2.5 72B Instruct
LMArena Longer Query14941282

Writing & Preference Kimi K3 leads

Kimi K3: 76.6 (#4), Qwen2.5 72B Instruct: 46.7 (#215)

Writing & Preference benchmarks
BenchmarkKimi K3Qwen2.5 72B Instruct
LMArena Text14761269
LMArena Creative Writing14541221
LMArena Multi-Turn14881272
EQ-Bench Creative Writing2082—
WildBench—80.2%
EQ-Bench 41339—

Frequently asked questions

Is Kimi K3 better than Qwen2.5 72B Instruct?

Kimi K3 is the stronger model overall, scoring 59.5 to 31.9 on the Noometry Index. Qwen2.5 72B Instruct costs 2.4× less per token, which makes it the better buy when Kimi K3's lead doesn't matter for your workload.

Which is cheaper, Kimi K3 or Qwen2.5 72B Instruct?

Qwen2.5 72B Instruct is cheaper. It lists at $1.40 per million input tokens and $5.60 per million output tokens; Kimi K3 lists at $3 and $15.

Is Kimi K3 or Qwen2.5 72B Instruct better for coding?

Kimi K3 scores higher on coding benchmarks: 61.0 versus 33.2 in the Noometry coding category.

Which has the bigger context window?

Kimi K3 does, with 1.05M tokens against 131K.

How many benchmarks do Kimi K3 and Qwen2.5 72B Instruct share?

24 benchmarks have published results for both models. Kimi K3 has 53 scored results on Noometry and Qwen2.5 72B Instruct has 43.

Related comparisons

Go deeper