Alibaba (Qwen), proprietary
Qwen2.5-Max
Qwen2.5-Max by Alibaba (Qwen) ranks 146th of 354 ranked models on the Noometry Index as of October 2026, with a score of 40.7. Its strongest category is coding, where it ranks 117th.
Last verified
Specifications
- Noometry rank
- #146 of 354
- Index score
- 40.7
- Evidence
- Confirmed 27 results
- Provider
Alibaba (Qwen)
- Released
- January 25, 2025
- Weights
- Proprietary
- Reasoning
- Unknown
- Context window
- —
- Max output
- —
- Input price
- Not listed
- Output price
- Not listed
- Blended price
- Not listed
- Output speed
- Not measured
- Value
- Not ranked
- Knowledge cutoff
- Unknown
Category scores
Each category score combines every public result we have in that category.
- Coding 41.8
- Reasoning 25.6
- Math 36.9
- Knowledge 35.3
- Multilingual 48.1
- Instruction Following 71.3
- Long Context 41.4
- Writing & Preference 55.4
| Category | Score | Rank | Results |
|---|---|---|---|
| Coding | 41.8 | #117 | 2 |
| Reasoning | 25.6 | #147 | 3 |
| Math | 36.9 | #162 | 2 |
| Knowledge | 35.3 | #186 | 2 |
| Multilingual | 48.1 | #146 | 1 |
| Instruction Following | 71.3 | #152 | 2 |
| Long Context | 41.4 | #142 | 1 |
| Writing & Preference | 55.4 | #146 | 5 |
Strengths and weaknesses
Categories where Qwen2.5-Max places highest and lowest among the models ranked in each, with its score against that category's median.
Strongest categories
| Category | Score | vs median | Rank |
|---|---|---|---|
| Coding | 41.8 | +3.1 | #117 of 340, top 35% |
| Reasoning | 25.6 | +1.9 | #147 of 350, top 42% |
| Writing & Preference | 55.4 | +1.6 | #146 of 312, top 47% |
Weakest categories
| Category | Score | vs median | Rank |
|---|---|---|---|
| Knowledge | 35.3 | −2.0 | #186 of 314, top 60% |
| Instruction Following | 71.3 | +0.0 | #152 of 305, top 50% |
| Math | 36.9 | +0.3 | #162 of 327, top 50% |
Closest competitors
The models ranked just above and below Qwen2.5-Max. When scores are this close, price and speed are often the better way to choose.
| Model | Rank | Score | Blended $/M | Speed | |
|---|---|---|---|---|---|
| Claude Opus 4.1 | #142 | 41.0 | $30 | — | Compare |
| o1 | #143 | 40.9 | $26.25 | — | Compare |
| Gemini 3.1 Flash Lite | #144 | 40.8 | $0.56 | 10 | Compare |
| Claude Sonnet 4 | #145 | 40.8 | $6 | 31 | Compare |
| Nemotron 3 Nano 30B A3B | #147 | 40.6 | $0.0875 | — | Compare |
| Granite 4.2 8B | #148 | 40.5 | $0.11 | — | Compare |
| Step 3 | #149 | 40.5 | — | 7 | Compare |
| MiniMax M1 | #150 | 40.3 | $0.96 | — | Compare |
Sponsored placements are available on pages like this one. Advertise on Noometry
Benchmark results
Every published result we track, with its source. Bold rows are the ones used for ranking; where several exist we prefer independent runs over self-reported numbers.
Coding
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LiveBench Coding | 64.4% | #11 of 39, top 29% | Epoch AI | ||
| LMArena Coding | 1359 | #168 of 294, top 58% | LMArena | 2026-10-08 |
Reasoning
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LiveBench Reasoning | 51.4% | #18 of 39, top 47% | Epoch AI | ||
| LMArena Hard Prompts | 1360 | #154 of 297, top 52% | LMArena | 2026-10-08 | |
| LiveBench Data Analysis | 67.9% | #8 of 39, top 21% | Epoch AI | ||
| Epoch Capabilities Index | 132.53 | #132 of 213, top 62% | Epoch AI | 2025-01-25 | |
| LiveBench | 62.3% | #12 of 39, top 31% | Epoch AI |
Math
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LiveBench Math | 58.4% | #14 of 39, top 36% | Epoch AI | ||
| LMArena Math | 1369 | #151 of 285, top 53% | LMArena | 2026-10-08 |
Knowledge
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| Confabulations (lower is better) | 21.8% | #36 of 51, top 71% | Lech Mazur benchmarks | ||
| LMArena Expert | 1337 | #163 of 273, top 60% | LMArena | 2026-10-08 |
Multilingual
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Non-English | 1352 | #146 of 297, top 50% | LMArena | 2026-10-08 | |
| LMArena Chinese | 1382 | #150 of 285, top 53% | LMArena | 2026-10-08 | |
| LMArena French | 1396 | #122 of 223, top 55% | LMArena | 2026-10-08 | |
| LMArena German | 1350 | #131 of 231, top 57% | LMArena | 2026-10-08 | |
| LMArena Japanese | 1300 | #123 of 211, top 59% | LMArena | 2026-10-08 | |
| LMArena Korean | 1304 | #130 of 213, top 62% | LMArena | 2026-10-08 | |
| LMArena Russian | 1353 | #146 of 283, top 52% | LMArena | 2026-10-08 | |
| LMArena Spanish | 1377 | #127 of 226, top 57% | LMArena | 2026-10-08 |
Instruction Following
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LiveBench Instruction Following | 75.3% | #14 of 39, top 36% | Epoch AI | ||
| LMArena Instruction Following | 1335 | #153 of 298, top 52% | LMArena | 2026-10-08 |
Long Context
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Longer Query | 1358 | #147 of 291, top 51% | LMArena | 2026-10-08 |
Writing & Preference
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Text | 1367 | #145 of 297, top 49% | LMArena | 2026-10-08 | |
| LMArena Creative Writing | 1339 | #140 of 295, top 48% | LMArena | 2026-10-08 | |
| Short-Story Creative Writing | 72.9% | #30 of 39, top 77% | Epoch AI | ||
| LMArena Multi-Turn | 1364 | #144 of 295, top 49% | LMArena | 2026-10-08 | |
| LiveBench Language | 56.3% | #6 of 39, top 16% | Epoch AI |
Compare Qwen2.5-Max
- Qwen2.5-Max vs Claude Sonnet 4
- Qwen2.5-Max vs Nemotron 3 Nano 30B A3B
- Qwen2.5-Max vs Gemini 3.1 Flash Lite
- Qwen2.5-Max vs Granite 4.2 8B
- Qwen2.5-Max vs o1
- Qwen2.5-Max vs Step 3
- Qwen2.5-Max vs GPT-6 Astra
- Qwen2.5-Max vs Claude Fable 5.1
- Qwen2.5-Max vs Gemini 3.8 Flash
- Qwen2.5-Max vs Kimi K3
- Qwen2.5-Max vs Grok 4.6
- Qwen2.5-Max vs GLM-5.3
- Qwen2.5-Max vs Muse Spark 1.3
- Qwen2.5-Max vs DeepSeek V4 Pro
Other Alibaba (Qwen) models
- Qwen3.8 Max56.8
- Qwen3.7 Max51.5
- Qwen3.6 Max Preview51.5
- Qwen3.6 Plus47.5
- Qwen3.5 397B-A17B46.0
- Qwen3.8 27B46.0
- Qwen3.5 Max Preview45.3
- Qwen3.7 Plus45.3
Frequently asked questions
How good is Qwen2.5-Max?
Qwen2.5-Max by Alibaba (Qwen) ranks 146th of 354 ranked models on the Noometry Index as of October 2026, with a score of 40.7. Its strongest category is coding, where it ranks 117th.
Is Qwen2.5-Max open source?
No. Qwen2.5-Max is proprietary and available only through Alibaba (Qwen)'s API and partner platforms.
What are Qwen2.5-Max's strengths and weaknesses?
Relative to other ranked models, Qwen2.5-Max places best in coding, reasoning, writing & preference and lowest in knowledge, instruction following, math.
What is Qwen2.5-Max best at?
Its best category is coding, where it ranks 117th on Noometry.