Alibaba (Qwen), open weights
Qwen2.5 72B Instruct
Qwen2.5 72B Instruct by Alibaba (Qwen) ranks 267th of 354 ranked models on the Noometry Index as of October 2026, with a score of 31.9. Its strongest category is agentic & tool use, where it ranks 133rd. API pricing starts at $1.40 per million input tokens and $5.60 per million output tokens, with a 131K-token context window.
Last verified
Specifications
- Noometry rank
- #267 of 354
- Index score
- 31.9
- Evidence
- Confirmed 43 results
- Provider
Alibaba (Qwen)
- Released
- September 1, 2024
- Weights
- Open weights
- Reasoning
- No
- Context window
- 131K
- Max output
- 8K
- Input price
- $1.40 / M
- Output price
- $5.60 / M
- Blended price
- $2.45 / M
- Output speed
- Not measured
- Value
- #172 of 219
- Knowledge cutoff
- April 2024
- Input
- text
- Hugging Face
- Qwen/Qwen2.5-72B-Instruct
Category scores
Each category score combines every public result we have in that category.
- Coding 33.2
- Agentic & Tool Use 22.1
- Reasoning 22.3
- Math 19.3
- Knowledge 27.0
- Multilingual 41.0
- Instruction Following 65.5
- Long Context 38.9
- Writing & Preference 46.7
| Category | Score | Rank | Results |
|---|---|---|---|
| Coding | 33.2 | #260 | 4 |
| Agentic & Tool Use | 22.1 | #133 | 2 |
| Reasoning | 22.3 | #199 | 3 |
| Math | 19.3 | #287 | 4 |
| Knowledge | 27.0 | #253 | 5 |
| Multilingual | 41.0 | #213 | 1 |
| Instruction Following | 65.5 | #221 | 2 |
| Long Context | 38.9 | #188 | 1 |
| Writing & Preference | 46.7 | #215 | 4 |
Strengths and weaknesses
Categories where Qwen2.5 72B Instruct places highest and lowest among the models ranked in each, with its score against that category's median.
Strongest categories
| Category | Score | vs median | Rank |
|---|---|---|---|
| Reasoning | 22.3 | −1.3 | #199 of 350, top 57% |
| Long Context | 38.9 | −2.0 | #188 of 296, top 64% |
| Writing & Preference | 46.7 | −7.1 | #215 of 312, top 69% |
Weakest categories
| Category | Score | vs median | Rank |
|---|---|---|---|
| Math | 19.3 | −17.3 | #287 of 327, top 88% |
| Agentic & Tool Use | 22.1 | −8.3 | #133 of 154, top 87% |
| Knowledge | 27.0 | −10.3 | #253 of 314, top 81% |
Closest competitors
The models ranked just above and below Qwen2.5 72B Instruct. When scores are this close, price and speed are often the better way to choose.
| Model | Rank | Score | Blended $/M | Speed | |
|---|---|---|---|---|---|
| Mistral Large | #263 | 31.9 | $3 | — | Compare |
| Qwen3-4B | #264 | 31.9 | — | — | Compare |
| Amazon Nova Lite | #265 | 31.9 | $0.10 | — | Compare |
| Mistral Medium 3.1 | #266 | 31.9 | $0.80 | — | Compare |
| Llama2 70b Steerlm Chat | #268 | 31.8 | — | — | Compare |
| Mistral Small 3.1 | #269 | 31.7 | $0.40 | — | Compare |
| Granite 3.0 8b Instruct | #270 | 31.6 | — | — | Compare |
| o1-pro | #271 | 31.5 | $263 | — | Compare |
Sponsored placements are available on pages like this one. Advertise on Noometry
Benchmark results
Every published result we track, with its source. Bold rows are the ones used for ranking; where several exist we prefer independent runs over self-reported numbers.
Coding
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| WeirdML | 16% | #108 of 119, top 91% | Epoch AI | ||
| BigCodeBench Instruct | 45.8% | #17 of 64, top 27% | BigCodeBench | 2024-09-19 | |
| LMArena Coding | 1292 | #202 of 294, top 69% | LMArena | 2026-10-08 | |
| BigCodeBench Complete | 55.9% | #16 of 66, top 25% | BigCodeBench | 2024-09-19 |
Agentic & Tool Use
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| TheAgentCompany | 5.7% | #11 of 14, top 79% | Epoch AI | ||
| BALROG | 16.2% | #29 of 35, top 83% | Epoch AI | ||
| METR Time Horizons | 35.8% | #29 of 32, top 91% | Epoch AI |
Reasoning
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Hard Prompts | 1271 | #205 of 297, top 70% | LMArena | 2026-10-08 | |
| DTBench | 62.9% | #106 of 151, top 71% | Epoch AI | ||
| LMCA | 13.4% | #107 of 125, top 86% | Epoch AI | ||
| BIG-Bench Hard | 79.8% | #5 of 27, top 19% | Epoch AI | ||
| Epoch Capabilities Index | 129 | #142 of 213, top 67% | Epoch AI | 2024-09-19 | |
| ForecastBench | 57.5 | #58 of 72, top 81% | Epoch AI | ||
| HellaSwag | 84.8% | #10 of 29, top 35% | Epoch AI | ||
| PIQA | 82.6% | #14 of 27, top 52% | Epoch AI | ||
| WinoGrande | 82.3% | #10 of 43, top 24% | Epoch AI |
Math
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| OTIS Mock AIME 2024-2025 | 8.1% | #136 of 173, top 79% | Epoch AI | 2025-02-25 | |
| Omni-MATH | 33% | #35 of 57, top 62% | HELM Capabilities | ||
| LMArena Math | 1283 | #191 of 285, top 68% | LMArena | 2026-10-08 | |
| MATH Level 5 | 63.2% | #37 of 79, top 47% | Epoch AI | 2025-01-27 |
Knowledge
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| GPQA Diamond | 49.1% | #129 of 186, top 70% | Epoch AI | 2025-01-27 | |
| MMLU-Pro | 63.1% | #40 of 58, top 69% | HELM Capabilities | ||
| Confabulations (lower is better) | 19.1% | #31 of 51, top 61% | Lech Mazur benchmarks | ||
| GPQA (HELM) | 42.6% | #41 of 57, top 72% | HELM Capabilities | ||
| LMArena Expert | 1245 | #201 of 273, top 74% | LMArena | 2026-10-08 | |
| ARC (AI2) Challenge | 94.5% | #3 of 39, top 8% | Epoch AI | ||
| MMLU | 85% | Epoch AI | |||
| MMLU | 85.3% | #7 of 81, top 9% | Epoch AI | ||
| TriviaQA | 71.9% | #19 of 25, top 76% | Epoch AI |
Multilingual
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Non-English | 1252 | #213 of 297, top 72% | LMArena | 2026-10-08 | |
| LMArena Chinese | 1272 | #197 of 285, top 70% | LMArena | 2026-10-08 | |
| LMArena French | 1280 | #171 of 223, top 77% | LMArena | 2026-10-08 | |
| LMArena German | 1234 | #182 of 231, top 79% | LMArena | 2026-10-08 | |
| LMArena Japanese | 1180 | #165 of 211, top 79% | LMArena | 2026-10-08 | |
| LMArena Korean | 1188 | #167 of 213, top 79% | LMArena | 2026-10-08 | |
| LMArena Russian | 1264 | #200 of 283, top 71% | LMArena | 2026-10-08 | |
| LMArena Spanish | 1256 | #181 of 226, top 81% | LMArena | 2026-10-08 |
Instruction Following
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| IFEval | 80.6% | #41 of 57, top 72% | HELM Capabilities | ||
| LMArena Instruction Following | 1254 | #207 of 298, top 70% | LMArena | 2026-10-08 |
Long Context
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Longer Query | 1282 | #202 of 291, top 70% | LMArena | 2026-10-08 |
Writing & Preference
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Text | 1269 | #217 of 297, top 74% | LMArena | 2026-10-08 | |
| LMArena Creative Writing | 1221 | #223 of 295, top 76% | LMArena | 2026-10-08 | |
| WildBench | 80.2% | #28 of 57, top 50% | HELM Capabilities | ||
| LMArena Multi-Turn | 1272 | #209 of 295, top 71% | LMArena | 2026-10-08 |
API pricing by provider
| Route | Input $/M | Output $/M | Cached input $/M | Checked |
|---|---|---|---|---|
| alibaba | $1.40 | $5.60 | — | 2026-10-10 |
| openrouter | $0.36 | $0.40 | — | 2026-10-10 |
Compare Qwen2.5 72B Instruct
- Qwen2.5 72B Instruct vs Mistral Medium 3.1
- Qwen2.5 72B Instruct vs Llama2 70b Steerlm Chat
- Qwen2.5 72B Instruct vs Amazon Nova Lite
- Qwen2.5 72B Instruct vs Mistral Small 3.1
- Qwen2.5 72B Instruct vs Qwen3-4B
- Qwen2.5 72B Instruct vs Granite 3.0 8b Instruct
- Qwen2.5 72B Instruct vs GPT-6 Astra
- Qwen2.5 72B Instruct vs Claude Fable 5.1
- Qwen2.5 72B Instruct vs Gemini 3.8 Flash
- Qwen2.5 72B Instruct vs Kimi K3
- Qwen2.5 72B Instruct vs Grok 4.6
- Qwen2.5 72B Instruct vs GLM-5.3
- Qwen2.5 72B Instruct vs Muse Spark 1.3
- Qwen2.5 72B Instruct vs DeepSeek V4 Pro
Other Alibaba (Qwen) models
- Qwen3.8 Max56.8
- Qwen3.7 Max51.5
- Qwen3.6 Max Preview51.5
- Qwen3.6 Plus47.5
- Qwen3.5 397B-A17B46.0
- Qwen3.8 27B46.0
- Qwen3.5 Max Preview45.3
- Qwen3.7 Plus45.3
Frequently asked questions
How good is Qwen2.5 72B Instruct?
Qwen2.5 72B Instruct by Alibaba (Qwen) ranks 267th of 354 ranked models on the Noometry Index as of October 2026, with a score of 31.9. Its strongest category is agentic & tool use, where it ranks 133rd. API pricing starts at $1.40 per million input tokens and $5.60 per million output tokens, with a 131K-token context window.
How much does Qwen2.5 72B Instruct cost?
Qwen2.5 72B Instruct costs $1.40 per million input tokens and $5.60 per million output tokens on Alibaba (Qwen)'s own API.
What is Qwen2.5 72B Instruct's context window?
Qwen2.5 72B Instruct accepts up to 131K tokens of input and can write up to 8K tokens in one response.
Is Qwen2.5 72B Instruct open source?
Yes. Qwen2.5 72B Instruct's weights are downloadable from Hugging Face (Qwen/Qwen2.5-72B-Instruct); check the license for commercial terms.
What are Qwen2.5 72B Instruct's strengths and weaknesses?
Relative to other ranked models, Qwen2.5 72B Instruct places best in reasoning, long context, writing & preference and lowest in math, agentic & tool use, knowledge.
What is Qwen2.5 72B Instruct best at?
Its best category is agentic & tool use, where it ranks 133rd on Noometry.