Alibaba (Qwen), proprietary

Qwen2.5-Max

Qwen2.5-Max by Alibaba (Qwen) ranks 146th of 354 ranked models on the Noometry Index as of October 2026, with a score of 40.7. Its strongest category is coding, where it ranks 117th.

Last verified

Specifications

Noometry rank
#146 of 354
Index score
40.7
Evidence
Confirmed 27 results
Released
January 25, 2025
Weights
Proprietary
Reasoning
Unknown
Context window
—
Max output
—
Input price
Not listed
Output price
Not listed
Blended price
Not listed
Output speed
Not measured
Value
Not ranked
Knowledge cutoff
Unknown

Category scores

Each category score combines every public result we have in that category.

Qwen2.5-Max category scores
  1. Coding 41.8
  2. Reasoning 25.6
  3. Math 36.9
  4. Knowledge 35.3
  5. Multilingual 48.1
  6. Instruction Following 71.3
  7. Long Context 41.4
  8. Writing & Preference 55.4
Qwen2.5-Max category ranks
CategoryScoreRankResults
Coding41.8#1172
Reasoning25.6#1473
Math36.9#1622
Knowledge35.3#1862
Multilingual48.1#1461
Instruction Following71.3#1522
Long Context41.4#1421
Writing & Preference55.4#1465

Strengths and weaknesses

Categories where Qwen2.5-Max places highest and lowest among the models ranked in each, with its score against that category's median.

Strongest categories

Qwen2.5-Max: strongest categories
CategoryScorevs medianRank
Coding41.8+3.1#117 of 340, top 35%
Reasoning25.6+1.9#147 of 350, top 42%
Writing & Preference55.4+1.6#146 of 312, top 47%

Weakest categories

Qwen2.5-Max: weakest categories
CategoryScorevs medianRank
Knowledge35.3−2.0#186 of 314, top 60%
Instruction Following71.3+0.0#152 of 305, top 50%
Math36.9+0.3#162 of 327, top 50%

Closest competitors

The models ranked just above and below Qwen2.5-Max. When scores are this close, price and speed are often the better way to choose.

Models ranked closest to Qwen2.5-Max
ModelRankScoreBlended $/MSpeed
Claude Opus 4.1#14241.0$30—Compare
o1#14340.9$26.25—Compare
Gemini 3.1 Flash Lite#14440.8$0.5610Compare
Claude Sonnet 4#14540.8$631Compare
Nemotron 3 Nano 30B A3B#14740.6$0.0875—Compare
Granite 4.2 8B#14840.5$0.11—Compare
Step 3#14940.5—7Compare
MiniMax M1#15040.3$0.96—Compare

Sponsored placements are available on pages like this one. Advertise on Noometry

Benchmark results

Every published result we track, with its source. Bold rows are the ones used for ranking; where several exist we prefer independent runs over self-reported numbers.

Coding

Qwen2.5-Max Coding benchmark results
BenchmarkScorePositionSettingSourceDate
LiveBench Coding64.4%#11 of 39, top 29%Epoch AI
LMArena Coding1359#168 of 294, top 58%LMArena2026-10-08

Reasoning

Qwen2.5-Max Reasoning benchmark results
BenchmarkScorePositionSettingSourceDate
LiveBench Reasoning51.4%#18 of 39, top 47%Epoch AI
LMArena Hard Prompts1360#154 of 297, top 52%LMArena2026-10-08
LiveBench Data Analysis67.9%#8 of 39, top 21%Epoch AI
Epoch Capabilities Index132.53#132 of 213, top 62%Epoch AI2025-01-25
LiveBench62.3%#12 of 39, top 31%Epoch AI

Math

Qwen2.5-Max Math benchmark results
BenchmarkScorePositionSettingSourceDate
LiveBench Math58.4%#14 of 39, top 36%Epoch AI
LMArena Math1369#151 of 285, top 53%LMArena2026-10-08

Knowledge

Qwen2.5-Max Knowledge benchmark results
BenchmarkScorePositionSettingSourceDate
Confabulations (lower is better)21.8%#36 of 51, top 71%Lech Mazur benchmarks
LMArena Expert1337#163 of 273, top 60%LMArena2026-10-08

Multilingual

Qwen2.5-Max Multilingual benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Non-English1352#146 of 297, top 50%LMArena2026-10-08
LMArena Chinese1382#150 of 285, top 53%LMArena2026-10-08
LMArena French1396#122 of 223, top 55%LMArena2026-10-08
LMArena German1350#131 of 231, top 57%LMArena2026-10-08
LMArena Japanese1300#123 of 211, top 59%LMArena2026-10-08
LMArena Korean1304#130 of 213, top 62%LMArena2026-10-08
LMArena Russian1353#146 of 283, top 52%LMArena2026-10-08
LMArena Spanish1377#127 of 226, top 57%LMArena2026-10-08

Instruction Following

Qwen2.5-Max Instruction Following benchmark results
BenchmarkScorePositionSettingSourceDate
LiveBench Instruction Following75.3%#14 of 39, top 36%Epoch AI
LMArena Instruction Following1335#153 of 298, top 52%LMArena2026-10-08

Long Context

Qwen2.5-Max Long Context benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Longer Query1358#147 of 291, top 51%LMArena2026-10-08

Writing & Preference

Qwen2.5-Max Writing & Preference benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Text1367#145 of 297, top 49%LMArena2026-10-08
LMArena Creative Writing1339#140 of 295, top 48%LMArena2026-10-08
Short-Story Creative Writing72.9%#30 of 39, top 77%Epoch AI
LMArena Multi-Turn1364#144 of 295, top 49%LMArena2026-10-08
LiveBench Language56.3%#6 of 39, top 16%Epoch AI

Compare Qwen2.5-Max

Other Alibaba (Qwen) models

Frequently asked questions

How good is Qwen2.5-Max?

Qwen2.5-Max by Alibaba (Qwen) ranks 146th of 354 ranked models on the Noometry Index as of October 2026, with a score of 40.7. Its strongest category is coding, where it ranks 117th.

Is Qwen2.5-Max open source?

No. Qwen2.5-Max is proprietary and available only through Alibaba (Qwen)'s API and partner platforms.

What are Qwen2.5-Max's strengths and weaknesses?

Relative to other ranked models, Qwen2.5-Max places best in coding, reasoning, writing & preference and lowest in knowledge, instruction following, math.

What is Qwen2.5-Max best at?

Its best category is coding, where it ranks 117th on Noometry.