Alibaba (Qwen), open weights

Qwen2.5 72B Instruct

Qwen2.5 72B Instruct by Alibaba (Qwen) ranks 267th of 354 ranked models on the Noometry Index as of October 2026, with a score of 31.9. Its strongest category is agentic & tool use, where it ranks 133rd. API pricing starts at $1.40 per million input tokens and $5.60 per million output tokens, with a 131K-token context window.

Last verified

Specifications

Noometry rank
#267 of 354
Index score
31.9
Evidence
Confirmed 43 results
Released
September 1, 2024
Weights
Open weights
Reasoning
No
Context window
131K
Max output
8K
Input price
$1.40 / M
Output price
$5.60 / M
Blended price
$2.45 / M
Output speed
Not measured
Value
#172 of 219
Knowledge cutoff
April 2024
Input
text

Category scores

Each category score combines every public result we have in that category.

Qwen2.5 72B Instruct category scores
  1. Coding 33.2
  2. Agentic & Tool Use 22.1
  3. Reasoning 22.3
  4. Math 19.3
  5. Knowledge 27.0
  6. Multilingual 41.0
  7. Instruction Following 65.5
  8. Long Context 38.9
  9. Writing & Preference 46.7
Qwen2.5 72B Instruct category ranks
CategoryScoreRankResults
Coding33.2#2604
Agentic & Tool Use22.1#1332
Reasoning22.3#1993
Math19.3#2874
Knowledge27.0#2535
Multilingual41.0#2131
Instruction Following65.5#2212
Long Context38.9#1881
Writing & Preference46.7#2154

Strengths and weaknesses

Categories where Qwen2.5 72B Instruct places highest and lowest among the models ranked in each, with its score against that category's median.

Strongest categories

Qwen2.5 72B Instruct: strongest categories
CategoryScorevs medianRank
Reasoning22.3−1.3#199 of 350, top 57%
Long Context38.9−2.0#188 of 296, top 64%
Writing & Preference46.7−7.1#215 of 312, top 69%

Weakest categories

Qwen2.5 72B Instruct: weakest categories
CategoryScorevs medianRank
Math19.3−17.3#287 of 327, top 88%
Agentic & Tool Use22.1−8.3#133 of 154, top 87%
Knowledge27.0−10.3#253 of 314, top 81%

Closest competitors

The models ranked just above and below Qwen2.5 72B Instruct. When scores are this close, price and speed are often the better way to choose.

Models ranked closest to Qwen2.5 72B Instruct
ModelRankScoreBlended $/MSpeed
Mistral Large#26331.9$3—Compare
Qwen3-4B#26431.9——Compare
Amazon Nova Lite#26531.9$0.10—Compare
Mistral Medium 3.1#26631.9$0.80—Compare
Llama2 70b Steerlm Chat#26831.8——Compare
Mistral Small 3.1#26931.7$0.40—Compare
Granite 3.0 8b Instruct#27031.6——Compare
o1-pro#27131.5$263—Compare

Sponsored placements are available on pages like this one. Advertise on Noometry

Benchmark results

Every published result we track, with its source. Bold rows are the ones used for ranking; where several exist we prefer independent runs over self-reported numbers.

Coding

Qwen2.5 72B Instruct Coding benchmark results
BenchmarkScorePositionSettingSourceDate
WeirdML16%#108 of 119, top 91%Epoch AI
BigCodeBench Instruct45.8%#17 of 64, top 27%BigCodeBench2024-09-19
LMArena Coding1292#202 of 294, top 69%LMArena2026-10-08
BigCodeBench Complete55.9%#16 of 66, top 25%BigCodeBench2024-09-19

Agentic & Tool Use

Qwen2.5 72B Instruct Agentic & Tool Use benchmark results
BenchmarkScorePositionSettingSourceDate
TheAgentCompany5.7%#11 of 14, top 79%Epoch AI
BALROG16.2%#29 of 35, top 83%Epoch AI
METR Time Horizons35.8%#29 of 32, top 91%Epoch AI

Reasoning

Qwen2.5 72B Instruct Reasoning benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Hard Prompts1271#205 of 297, top 70%LMArena2026-10-08
DTBench62.9%#106 of 151, top 71%Epoch AI
LMCA13.4%#107 of 125, top 86%Epoch AI
BIG-Bench Hard79.8%#5 of 27, top 19%Epoch AI
Epoch Capabilities Index129#142 of 213, top 67%Epoch AI2024-09-19
ForecastBench57.5#58 of 72, top 81%Epoch AI
HellaSwag84.8%#10 of 29, top 35%Epoch AI
PIQA82.6%#14 of 27, top 52%Epoch AI
WinoGrande82.3%#10 of 43, top 24%Epoch AI

Math

Qwen2.5 72B Instruct Math benchmark results
BenchmarkScorePositionSettingSourceDate
OTIS Mock AIME 2024-20258.1%#136 of 173, top 79%Epoch AI2025-02-25
Omni-MATH33%#35 of 57, top 62%HELM Capabilities
LMArena Math1283#191 of 285, top 68%LMArena2026-10-08
MATH Level 563.2%#37 of 79, top 47%Epoch AI2025-01-27

Knowledge

Qwen2.5 72B Instruct Knowledge benchmark results
BenchmarkScorePositionSettingSourceDate
GPQA Diamond49.1%#129 of 186, top 70%Epoch AI2025-01-27
MMLU-Pro63.1%#40 of 58, top 69%HELM Capabilities
Confabulations (lower is better)19.1%#31 of 51, top 61%Lech Mazur benchmarks
GPQA (HELM)42.6%#41 of 57, top 72%HELM Capabilities
LMArena Expert1245#201 of 273, top 74%LMArena2026-10-08
ARC (AI2) Challenge94.5%#3 of 39, top 8%Epoch AI
MMLU85%Epoch AI
MMLU85.3%#7 of 81, top 9%Epoch AI
TriviaQA71.9%#19 of 25, top 76%Epoch AI

Multilingual

Qwen2.5 72B Instruct Multilingual benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Non-English1252#213 of 297, top 72%LMArena2026-10-08
LMArena Chinese1272#197 of 285, top 70%LMArena2026-10-08
LMArena French1280#171 of 223, top 77%LMArena2026-10-08
LMArena German1234#182 of 231, top 79%LMArena2026-10-08
LMArena Japanese1180#165 of 211, top 79%LMArena2026-10-08
LMArena Korean1188#167 of 213, top 79%LMArena2026-10-08
LMArena Russian1264#200 of 283, top 71%LMArena2026-10-08
LMArena Spanish1256#181 of 226, top 81%LMArena2026-10-08

Instruction Following

Qwen2.5 72B Instruct Instruction Following benchmark results
BenchmarkScorePositionSettingSourceDate
IFEval80.6%#41 of 57, top 72%HELM Capabilities
LMArena Instruction Following1254#207 of 298, top 70%LMArena2026-10-08

Long Context

Qwen2.5 72B Instruct Long Context benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Longer Query1282#202 of 291, top 70%LMArena2026-10-08

Writing & Preference

Qwen2.5 72B Instruct Writing & Preference benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Text1269#217 of 297, top 74%LMArena2026-10-08
LMArena Creative Writing1221#223 of 295, top 76%LMArena2026-10-08
WildBench80.2%#28 of 57, top 50%HELM Capabilities
LMArena Multi-Turn1272#209 of 295, top 71%LMArena2026-10-08

API pricing by provider

Qwen2.5 72B Instruct API prices
RouteInput $/MOutput $/MCached input $/MChecked
alibaba$1.40$5.60—2026-10-10
openrouter$0.36$0.40—2026-10-10

Compare Qwen2.5 72B Instruct

Other Alibaba (Qwen) models

Frequently asked questions

How good is Qwen2.5 72B Instruct?

Qwen2.5 72B Instruct by Alibaba (Qwen) ranks 267th of 354 ranked models on the Noometry Index as of October 2026, with a score of 31.9. Its strongest category is agentic & tool use, where it ranks 133rd. API pricing starts at $1.40 per million input tokens and $5.60 per million output tokens, with a 131K-token context window.

How much does Qwen2.5 72B Instruct cost?

Qwen2.5 72B Instruct costs $1.40 per million input tokens and $5.60 per million output tokens on Alibaba (Qwen)'s own API.

What is Qwen2.5 72B Instruct's context window?

Qwen2.5 72B Instruct accepts up to 131K tokens of input and can write up to 8K tokens in one response.

Is Qwen2.5 72B Instruct open source?

Yes. Qwen2.5 72B Instruct's weights are downloadable from Hugging Face (Qwen/Qwen2.5-72B-Instruct); check the license for commercial terms.

What are Qwen2.5 72B Instruct's strengths and weaknesses?

Relative to other ranked models, Qwen2.5 72B Instruct places best in reasoning, long context, writing & preference and lowest in math, agentic & tool use, knowledge.

What is Qwen2.5 72B Instruct best at?

Its best category is agentic & tool use, where it ranks 133rd on Noometry.