Alibaba (Qwen), open weights

# Qwen2.5 72B Instruct

> Qwen2.5 72B Instruct by Alibaba (Qwen), released September 2024. Ranked #267 of 354 with a Noometry Index of 31.9. API: $1.40 in / $5.60 out per M tokens. 131K context. Scores, sources and comparisons.
- Canonical page: https://noometry.com/models/qwen2-5-72b-instruct
- Last updated: 2026-10-10
- Title: Qwen2.5 72B Instruct Benchmarks, Price & Rank (October 2026)

Qwen2.5 72B Instruct by Alibaba (Qwen) ranks 267th of 354 ranked models on the Noometry Index as of October 2026, with a score of 31.9. Its strongest category is agentic & tool use, where it ranks 133rd. API pricing starts at $1.40 per million input tokens and $5.60 per million output tokens, with a 131K-token context window.

Last verified October 10, 2026

## Specifications

- **Noometry rank:** #267 of 354
- **Index score:** 31.9
- **Evidence:** Confirmed 43 results
- **Provider:** [![](/logos/alibaba.svg) Alibaba (Qwen)](https://noometry.com/providers/alibaba)
- **Released:** September 1, 2024
- **Weights:** Open weights
- **Reasoning:** No
- **Context window:** 131K
- **Max output:** 8K
- **Input price:** $1.40 / M
- **Output price:** $5.60 / M
- **Blended price:** $2.45 / M
- **Output speed:** Not measured
- **Value:** #172 of 219
- **Knowledge cutoff:** April 2024
- **Input:** text
- **Hugging Face:** [Qwen/Qwen2.5-72B-Instruct](https://huggingface.co/Qwen/Qwen2.5-72B-Instruct)

## Category scores

Each category score combines every public result we have in that category.

Qwen2.5 72B Instruct category scores

1.  Coding 33.2
2.  Agentic & Tool Use 22.1
3.  Reasoning 22.3
4.  Math 19.3
5.  Knowledge 27.0
6.  Multilingual 41.0
7.  Instruction Following 65.5
8.  Long Context 38.9
9.  Writing & Preference 46.7
10.  020406080

Qwen2.5 72B Instruct category ranks
| Category | Score | Rank | Results |
| --- | --- | --- | --- |
| [Coding](https://noometry.com/best/coding) | 33.2 | #260 | 4 |
| [Agentic & Tool Use](https://noometry.com/best/agentic) | 22.1 | #133 | 2 |
| [Reasoning](https://noometry.com/best/reasoning) | 22.3 | #199 | 3 |
| [Math](https://noometry.com/best/math) | 19.3 | #287 | 4 |
| [Knowledge](https://noometry.com/best/knowledge) | 27.0 | #253 | 5 |
| [Multilingual](https://noometry.com/best/multilingual) | 41.0 | #213 | 1 |
| [Instruction Following](https://noometry.com/best/instruction-following) | 65.5 | #221 | 2 |
| [Long Context](https://noometry.com/best/long-context) | 38.9 | #188 | 1 |
| [Writing & Preference](https://noometry.com/best/writing) | 46.7 | #215 | 4 |

## Strengths and weaknesses

Categories where Qwen2.5 72B Instruct places highest and lowest among the models ranked in each, with its score against that category's median.

### Strongest categories

Qwen2.5 72B Instruct: strongest categories
| Category | Score | vs median | Rank |
| --- | --- | --- | --- |
| [Reasoning](https://noometry.com/best/reasoning) | 22.3 | −1.3 | #199 of 350, top 57% |
| [Long Context](https://noometry.com/best/long-context) | 38.9 | −2.0 | #188 of 296, top 64% |
| [Writing & Preference](https://noometry.com/best/writing) | 46.7 | −7.1 | #215 of 312, top 69% |

### Weakest categories

Qwen2.5 72B Instruct: weakest categories
| Category | Score | vs median | Rank |
| --- | --- | --- | --- |
| [Math](https://noometry.com/best/math) | 19.3 | −17.3 | #287 of 327, top 88% |
| [Agentic & Tool Use](https://noometry.com/best/agentic) | 22.1 | −8.3 | #133 of 154, top 87% |
| [Knowledge](https://noometry.com/best/knowledge) | 27.0 | −10.3 | #253 of 314, top 81% |

## Closest competitors

The models ranked just above and below Qwen2.5 72B Instruct. When scores are this close, price and speed are often the better way to choose.

Models ranked closest to Qwen2.5 72B Instruct
| Model | Rank | Score | Blended $/M | Speed |  |
| --- | --- | --- | --- | --- | --- |
| [Mistral Large](https://noometry.com/models/mistral-large) | #263 | 31.9 | $3 | — | [Compare](https://noometry.com/compare/mistral-large-vs-qwen2-5-72b-instruct) |
| [Qwen3-4B](https://noometry.com/models/qwen3-4b) | #264 | 31.9 | — | — | [Compare](https://noometry.com/compare/qwen2-5-72b-instruct-vs-qwen3-4b) |
| [Amazon Nova Lite](https://noometry.com/models/amazon-nova-lite) | #265 | 31.9 | $0.10 | — | [Compare](https://noometry.com/compare/amazon-nova-lite-vs-qwen2-5-72b-instruct) |
| [Mistral Medium 3.1](https://noometry.com/models/mistral-medium-3-1) | #266 | 31.9 | $0.80 | — | [Compare](https://noometry.com/compare/mistral-medium-3-1-vs-qwen2-5-72b-instruct) |
| [Llama2 70b Steerlm Chat](https://noometry.com/models/llama2-70b-steerlm-chat) | #268 | 31.8 | — | — | [Compare](https://noometry.com/compare/llama2-70b-steerlm-chat-vs-qwen2-5-72b-instruct) |
| [Mistral Small 3.1](https://noometry.com/models/mistral-small-3-1) | #269 | 31.7 | $0.40 | — | [Compare](https://noometry.com/compare/mistral-small-3-1-vs-qwen2-5-72b-instruct) |
| [Granite 3.0 8b Instruct](https://noometry.com/models/granite-3-0-8b-instruct) | #270 | 31.6 | — | — | [Compare](https://noometry.com/compare/granite-3-0-8b-instruct-vs-qwen2-5-72b-instruct) |
| [o1-pro](https://noometry.com/models/o1-pro) | #271 | 31.5 | $263 | — | [Compare](https://noometry.com/compare/o1-pro-vs-qwen2-5-72b-instruct) |

Sponsored placements are available on pages like this one. [Advertise on Noometry](https://noometry.com/advertise)

## Benchmark results

Every published result we track, with its source. Bold rows are the ones used for ranking; where several exist we prefer independent runs over self-reported numbers.

### Coding

Qwen2.5 72B Instruct Coding benchmark results
| Benchmark | Score | Position | Setting | Source | Date |
| --- | --- | --- | --- | --- | --- |
| [WeirdML](https://noometry.com/benchmarks/weirdml) | 16% | #108 of 119, top 91% |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| [BigCodeBench Instruct](https://noometry.com/benchmarks/bigcodebench-instruct) | 45.8% | #17 of 64, top 27% |  | [BigCodeBench](https://bigcode-bench.github.io/) | 2024-09-19 |
| [LMArena Coding](https://noometry.com/benchmarks/arena-coding) | 1292 | #202 of 294, top 69% |  | [LMArena](https://lmarena.ai/leaderboard/text) | 2026-10-08 |
| [BigCodeBench Complete](https://noometry.com/benchmarks/bigcodebench-complete) | 55.9% | #16 of 66, top 25% |  | [BigCodeBench](https://bigcode-bench.github.io/) | 2024-09-19 |

### Agentic & Tool Use

Qwen2.5 72B Instruct Agentic & Tool Use benchmark results
| Benchmark | Score | Position | Setting | Source | Date |
| --- | --- | --- | --- | --- | --- |
| [TheAgentCompany](https://noometry.com/benchmarks/the-agent-company) | 5.7% | #11 of 14, top 79% |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| [BALROG](https://noometry.com/benchmarks/balrog) | 16.2% | #29 of 35, top 83% |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| [METR Time Horizons](https://noometry.com/benchmarks/metr-time-horizons) | 35.8% | #29 of 32, top 91% |  | [Epoch AI](https://epoch.ai/benchmarks) |  |

### Reasoning

Qwen2.5 72B Instruct Reasoning benchmark results
| Benchmark | Score | Position | Setting | Source | Date |
| --- | --- | --- | --- | --- | --- |
| [LMArena Hard Prompts](https://noometry.com/benchmarks/arena-hard-prompts) | 1271 | #205 of 297, top 70% |  | [LMArena](https://lmarena.ai/leaderboard/text) | 2026-10-08 |
| [DTBench](https://noometry.com/benchmarks/dtbench) | 62.9% | #106 of 151, top 71% |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| [LMCA](https://noometry.com/benchmarks/lmca) | 13.4% | #107 of 125, top 86% |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| [BIG-Bench Hard](https://noometry.com/benchmarks/bbh) | 79.8% | #5 of 27, top 19% |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| [Epoch Capabilities Index](https://noometry.com/benchmarks/epoch-capabilities-index) | 129 | #142 of 213, top 67% |  | [Epoch AI](https://epoch.ai/eci) | 2024-09-19 |
| [ForecastBench](https://noometry.com/benchmarks/forecastbench) | 57.5 | #58 of 72, top 81% |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| [HellaSwag](https://noometry.com/benchmarks/hellaswag) | 84.8% | #10 of 29, top 35% |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| [PIQA](https://noometry.com/benchmarks/piqa) | 82.6% | #14 of 27, top 52% |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| [WinoGrande](https://noometry.com/benchmarks/winogrande) | 82.3% | #10 of 43, top 24% |  | [Epoch AI](https://epoch.ai/benchmarks) |  |

### Math

Qwen2.5 72B Instruct Math benchmark results
| Benchmark | Score | Position | Setting | Source | Date |
| --- | --- | --- | --- | --- | --- |
| [OTIS Mock AIME 2024-2025](https://noometry.com/benchmarks/otis-mock-aime) | 8.1% | #136 of 173, top 79% |  | [Epoch AI](https://epoch.ai/benchmarks) | 2025-02-25 |
| [Omni-MATH](https://noometry.com/benchmarks/omni-math) | 33% | #35 of 57, top 62% |  | [HELM Capabilities](https://crfm.stanford.edu/helm/capabilities/latest/) |  |
| [LMArena Math](https://noometry.com/benchmarks/arena-math) | 1283 | #191 of 285, top 68% |  | [LMArena](https://lmarena.ai/leaderboard/text) | 2026-10-08 |
| [MATH Level 5](https://noometry.com/benchmarks/math-level-5) | 63.2% | #37 of 79, top 47% |  | [Epoch AI](https://epoch.ai/benchmarks) | 2025-01-27 |

### Knowledge

Qwen2.5 72B Instruct Knowledge benchmark results
| Benchmark | Score | Position | Setting | Source | Date |
| --- | --- | --- | --- | --- | --- |
| [GPQA Diamond](https://noometry.com/benchmarks/gpqa-diamond) | 49.1% | #129 of 186, top 70% |  | [Epoch AI](https://epoch.ai/benchmarks) | 2025-01-27 |
| [MMLU-Pro](https://noometry.com/benchmarks/mmlu-pro) | 63.1% | #40 of 58, top 69% |  | [HELM Capabilities](https://crfm.stanford.edu/helm/capabilities/latest/) |  |
| [Confabulations](https://noometry.com/benchmarks/confabulations) (lower is better) | 19.1% | #31 of 51, top 61% |  | [Lech Mazur benchmarks](https://github.com/lechmazur/confabulations) |  |
| [GPQA (HELM)](https://noometry.com/benchmarks/helm-gpqa) | 42.6% | #41 of 57, top 72% |  | [HELM Capabilities](https://crfm.stanford.edu/helm/capabilities/latest/) |  |
| [LMArena Expert](https://noometry.com/benchmarks/arena-expert) | 1245 | #201 of 273, top 74% |  | [LMArena](https://lmarena.ai/leaderboard/text) | 2026-10-08 |
| [ARC (AI2) Challenge](https://noometry.com/benchmarks/arc-challenge) | 94.5% | #3 of 39, top 8% |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| [MMLU](https://noometry.com/benchmarks/mmlu) | 85% |  |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| [MMLU](https://noometry.com/benchmarks/mmlu) | 85.3% | #7 of 81, top 9% |  | [Epoch AI](https://epoch.ai/benchmarks) |  |
| [TriviaQA](https://noometry.com/benchmarks/triviaqa) | 71.9% | #19 of 25, top 76% |  | [Epoch AI](https://epoch.ai/benchmarks) |  |

### Multilingual

Qwen2.5 72B Instruct Multilingual benchmark results
| Benchmark | Score | Position | Setting | Source | Date |
| --- | --- | --- | --- | --- | --- |
| [LMArena Non-English](https://noometry.com/benchmarks/arena-non-english) | 1252 | #213 of 297, top 72% |  | [LMArena](https://lmarena.ai/leaderboard/text) | 2026-10-08 |
| [LMArena Chinese](https://noometry.com/benchmarks/arena-chinese) | 1272 | #197 of 285, top 70% |  | [LMArena](https://lmarena.ai/leaderboard/text) | 2026-10-08 |
| [LMArena French](https://noometry.com/benchmarks/arena-french) | 1280 | #171 of 223, top 77% |  | [LMArena](https://lmarena.ai/leaderboard/text) | 2026-10-08 |
| [LMArena German](https://noometry.com/benchmarks/arena-german) | 1234 | #182 of 231, top 79% |  | [LMArena](https://lmarena.ai/leaderboard/text) | 2026-10-08 |
| [LMArena Japanese](https://noometry.com/benchmarks/arena-japanese) | 1180 | #165 of 211, top 79% |  | [LMArena](https://lmarena.ai/leaderboard/text) | 2026-10-08 |
| [LMArena Korean](https://noometry.com/benchmarks/arena-korean) | 1188 | #167 of 213, top 79% |  | [LMArena](https://lmarena.ai/leaderboard/text) | 2026-10-08 |
| [LMArena Russian](https://noometry.com/benchmarks/arena-russian) | 1264 | #200 of 283, top 71% |  | [LMArena](https://lmarena.ai/leaderboard/text) | 2026-10-08 |
| [LMArena Spanish](https://noometry.com/benchmarks/arena-spanish) | 1256 | #181 of 226, top 81% |  | [LMArena](https://lmarena.ai/leaderboard/text) | 2026-10-08 |

### Instruction Following

Qwen2.5 72B Instruct Instruction Following benchmark results
| Benchmark | Score | Position | Setting | Source | Date |
| --- | --- | --- | --- | --- | --- |
| [IFEval](https://noometry.com/benchmarks/ifeval) | 80.6% | #41 of 57, top 72% |  | [HELM Capabilities](https://crfm.stanford.edu/helm/capabilities/latest/) |  |
| [LMArena Instruction Following](https://noometry.com/benchmarks/arena-instruction-following) | 1254 | #207 of 298, top 70% |  | [LMArena](https://lmarena.ai/leaderboard/text) | 2026-10-08 |

### Long Context

Qwen2.5 72B Instruct Long Context benchmark results
| Benchmark | Score | Position | Setting | Source | Date |
| --- | --- | --- | --- | --- | --- |
| [LMArena Longer Query](https://noometry.com/benchmarks/arena-longer-query) | 1282 | #202 of 291, top 70% |  | [LMArena](https://lmarena.ai/leaderboard/text) | 2026-10-08 |

### Writing & Preference

Qwen2.5 72B Instruct Writing & Preference benchmark results
| Benchmark | Score | Position | Setting | Source | Date |
| --- | --- | --- | --- | --- | --- |
| [LMArena Text](https://noometry.com/benchmarks/arena-text) | 1269 | #217 of 297, top 74% |  | [LMArena](https://lmarena.ai/leaderboard/text) | 2026-10-08 |
| [LMArena Creative Writing](https://noometry.com/benchmarks/arena-creative-writing) | 1221 | #223 of 295, top 76% |  | [LMArena](https://lmarena.ai/leaderboard/text) | 2026-10-08 |
| [WildBench](https://noometry.com/benchmarks/wildbench) | 80.2% | #28 of 57, top 50% |  | [HELM Capabilities](https://crfm.stanford.edu/helm/capabilities/latest/) |  |
| [LMArena Multi-Turn](https://noometry.com/benchmarks/arena-multi-turn) | 1272 | #209 of 295, top 71% |  | [LMArena](https://lmarena.ai/leaderboard/text) | 2026-10-08 |

## API pricing by provider

Qwen2.5 72B Instruct API prices
| Route | Input $/M | Output $/M | Cached input $/M | Checked |
| --- | --- | --- | --- | --- |
| [alibaba](https://www.alibabacloud.com/help/en/model-studio/models) | $1.40 | $5.60 | — | 2026-10-10 |
| [openrouter](https://openrouter.ai/qwen/qwen-2.5-72b-instruct) | $0.36 | $0.40 | — | 2026-10-10 |

[All Alibaba (Qwen) API prices →](https://noometry.com/llm-pricing/alibaba) [Estimate your cost →](https://noometry.com/tools/cost-calculator)

## Compare Qwen2.5 72B Instruct

-   [Qwen2.5 72B Instruct vs Mistral Medium 3.1](https://noometry.com/compare/mistral-medium-3-1-vs-qwen2-5-72b-instruct)
-   [Qwen2.5 72B Instruct vs Llama2 70b Steerlm Chat](https://noometry.com/compare/llama2-70b-steerlm-chat-vs-qwen2-5-72b-instruct)
-   [Qwen2.5 72B Instruct vs Amazon Nova Lite](https://noometry.com/compare/amazon-nova-lite-vs-qwen2-5-72b-instruct)
-   [Qwen2.5 72B Instruct vs Mistral Small 3.1](https://noometry.com/compare/mistral-small-3-1-vs-qwen2-5-72b-instruct)
-   [Qwen2.5 72B Instruct vs Qwen3-4B](https://noometry.com/compare/qwen2-5-72b-instruct-vs-qwen3-4b)
-   [Qwen2.5 72B Instruct vs Granite 3.0 8b Instruct](https://noometry.com/compare/granite-3-0-8b-instruct-vs-qwen2-5-72b-instruct)
-   [Qwen2.5 72B Instruct vs GPT-6 Astra](https://noometry.com/compare/gpt-6-astra-vs-qwen2-5-72b-instruct)
-   [Qwen2.5 72B Instruct vs Claude Fable 5.1](https://noometry.com/compare/claude-fable-5-1-vs-qwen2-5-72b-instruct)
-   [Qwen2.5 72B Instruct vs Gemini 3.8 Flash](https://noometry.com/compare/gemini-3-8-flash-vs-qwen2-5-72b-instruct)
-   [Qwen2.5 72B Instruct vs Kimi K3](https://noometry.com/compare/kimi-k3-vs-qwen2-5-72b-instruct)
-   [Qwen2.5 72B Instruct vs Grok 4.6](https://noometry.com/compare/grok-4-6-vs-qwen2-5-72b-instruct)
-   [Qwen2.5 72B Instruct vs GLM-5.3](https://noometry.com/compare/glm-5-3-vs-qwen2-5-72b-instruct)
-   [Qwen2.5 72B Instruct vs Muse Spark 1.3](https://noometry.com/compare/muse-spark-1-3-vs-qwen2-5-72b-instruct)
-   [Qwen2.5 72B Instruct vs DeepSeek V4 Pro](https://noometry.com/compare/deepseek-v4-pro-vs-qwen2-5-72b-instruct)

## Other Alibaba (Qwen) models

-   [Qwen3.8 Max](https://noometry.com/models/qwen3-8-max)56.8
-   [Qwen3.7 Max](https://noometry.com/models/qwen3-7-max)51.5
-   [Qwen3.6 Max Preview](https://noometry.com/models/qwen3-6-max-preview)51.5
-   [Qwen3.6 Plus](https://noometry.com/models/qwen3-6-plus)47.5
-   [Qwen3.5 397B-A17B](https://noometry.com/models/qwen3-5-397b-a17b)46.0
-   [Qwen3.8 27B](https://noometry.com/models/qwen3-8-27b)46.0
-   [Qwen3.5 Max Preview](https://noometry.com/models/qwen3-5-max-preview)45.3
-   [Qwen3.7 Plus](https://noometry.com/models/qwen3-7-plus)45.3

## Frequently asked questions

### How good is Qwen2.5 72B Instruct?

Qwen2.5 72B Instruct by Alibaba (Qwen) ranks 267th of 354 ranked models on the Noometry Index as of October 2026, with a score of 31.9. Its strongest category is agentic & tool use, where it ranks 133rd. API pricing starts at $1.40 per million input tokens and $5.60 per million output tokens, with a 131K-token context window.

### How much does Qwen2.5 72B Instruct cost?

Qwen2.5 72B Instruct costs $1.40 per million input tokens and $5.60 per million output tokens on Alibaba (Qwen)'s own API.

### What is Qwen2.5 72B Instruct's context window?

Qwen2.5 72B Instruct accepts up to 131K tokens of input and can write up to 8K tokens in one response.

### Is Qwen2.5 72B Instruct open source?

Yes. Qwen2.5 72B Instruct's weights are downloadable from Hugging Face (Qwen/Qwen2.5-72B-Instruct); check the license for commercial terms.

### What are Qwen2.5 72B Instruct's strengths and weaknesses?

Relative to other ranked models, Qwen2.5 72B Instruct places best in reasoning, long context, writing & preference and lowest in math, agentic & tool use, knowledge.

### What is Qwen2.5 72B Instruct best at?

Its best category is agentic & tool use, where it ranks 133rd on Noometry.

### Cite this page

Noometry. (2026). Qwen2.5 72B Instruct benchmarks and pricing. Retrieved October 10, 2026, from https://noometry.com/models/qwen2-5-72b-instruct

Quote Noometry with a link back to this page. It is also available in [Markdown](https://noometry.com/md/models/qwen2-5-72b-instruct.md).
