Model comparison

Qwen Max vs Qwen2.5 7B Instruct

Qwen Max is the stronger model overall, scoring 34.7 to 29.0 on the Noometry Index. Qwen2.5 7B Instruct costs 9.1× less per token, which makes it the better buy when Qwen Max's lead doesn't matter for your workload.

Last verified . 2 shared benchmarks.

Qwen Max Alibaba (Qwen)

34.7

Rank #230 Confirmed

Qwen2.5 7B Instruct Alibaba (Qwen)

29.0

Rank #320 Confirmed

Summary

  • They share 2 benchmarks with published results for both. Qwen Max scores higher in 4 categories and Qwen2.5 7B Instruct in 2 categories; 5 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Qwen Max leads 30.3 to 17.0.
  • The biggest single-benchmark swing is GPQA Diamond: 56.1% for Qwen Max and 35.5% for Qwen2.5 7B Instruct.
  • Qwen2.5 7B Instruct is cheaper at $0.17 / $0.70 per million input/output tokens, against $1.60 / $6.40 for Qwen Max.
  • Qwen2.5 7B Instruct accepts more context: 131K tokens versus 33K.
  • Qwen2.5 7B Instruct has downloadable open weights; the other is API-only.

Side by side

Qwen Max and Qwen2.5 7B Instruct specifications
Qwen MaxQwen2.5 7B Instruct
ProviderAlibaba (Qwen)Alibaba (Qwen)
Noometry Index34.729.0
Released2024-04-032024-09
WeightsProprietaryOpen
Context window33K131K
Max output8K8K
Input $ / M tokens$1.60$0.17
Output $ / M tokens$6.40$0.70
Results tracked2315

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Qwen2.5 7B Instruct leads

Qwen Max: 30.7 (#292), Qwen2.5 7B Instruct: 36.5 (#208)

Coding benchmarks
BenchmarkQwen MaxQwen2.5 7B Instruct
Aider Polyglot21.8%—
BigCodeBench Instruct—37.6%
LMArena Coding1288—
BigCodeBench Complete—46.1%

Agentic & Tool Use Not comparable

Qwen Max: —, Qwen2.5 7B Instruct: 23.8 (#124)

Agentic & Tool Use benchmarks
BenchmarkQwen MaxQwen2.5 7B Instruct
BALROG—7.8%

Reasoning Qwen Max leads

Qwen Max: 25.1 (#151), Qwen2.5 7B Instruct: 14.8 (#322)

Reasoning benchmarks
BenchmarkQwen MaxQwen2.5 7B Instruct
Chess Puzzles—0%
LMArena Hard Prompts1269—
DTBench—47.7%
LMCA—6.4%
Epoch Capabilities Index—118.51

Math Qwen Max leads

Qwen Max: 22.3 (#276), Qwen2.5 7B Instruct: 12.6 (#306)

Math benchmarks
BenchmarkQwen MaxQwen2.5 7B Instruct
OTIS Mock AIME 2024-202516.1%2.5%
Omni-MATH—29.4%
LMArena Math1275—
MATH Level 567.2%—
FrontierMath (Feb 2025 set)1%—

Knowledge Qwen Max leads

Qwen Max: 30.3 (#228), Qwen2.5 7B Instruct: 17.0 (#286)

Knowledge benchmarks
BenchmarkQwen MaxQwen2.5 7B Instruct
GPQA Diamond56.1%35.5%
MMLU-Pro—53.9%
GPQA (HELM)—34.1%
LMArena Expert1248—
MMLU—72.9%

Multilingual Not comparable

Qwen Max: 41.8 (#202), Qwen2.5 7B Instruct: —

Multilingual benchmarks
BenchmarkQwen MaxQwen2.5 7B Instruct
LMArena Non-English1263—
LMArena Chinese1254—
LMArena French1330—
LMArena German1254—
LMArena Japanese1205—
LMArena Korean1142—
LMArena Russian1274—
LMArena Spanish1290—

Instruction Following Qwen Max leads

Qwen Max: 66.5 (#208), Qwen2.5 7B Instruct: 63.2 (#231)

Instruction Following benchmarks
BenchmarkQwen MaxQwen2.5 7B Instruct
IFEval—74.1%
LMArena Instruction Following1262—

Long Context Not comparable

Qwen Max: 39.4 (#180), Qwen2.5 7B Instruct: —

Long Context benchmarks
BenchmarkQwen MaxQwen2.5 7B Instruct
Fiction.LiveBench66.7%—
LMArena Longer Query1288—

Writing & Preference Too close to call

Qwen Max: 47.8 (#205), Qwen2.5 7B Instruct: 48.8 (#195)

Writing & Preference benchmarks
BenchmarkQwen MaxQwen2.5 7B Instruct
LMArena Text1282—
LMArena Creative Writing1248—
WildBench—73.1%
LMArena Multi-Turn1277—

Frequently asked questions

Is Qwen Max better than Qwen2.5 7B Instruct?

Qwen Max is the stronger model overall, scoring 34.7 to 29.0 on the Noometry Index. Qwen2.5 7B Instruct costs 9.1× less per token, which makes it the better buy when Qwen Max's lead doesn't matter for your workload.

Which is cheaper, Qwen Max or Qwen2.5 7B Instruct?

Qwen2.5 7B Instruct is cheaper. It lists at $0.17 per million input tokens and $0.70 per million output tokens; Qwen Max lists at $1.60 and $6.40.

Is Qwen Max or Qwen2.5 7B Instruct better for coding?

Qwen2.5 7B Instruct scores higher on coding benchmarks: 36.5 versus 30.7 in the Noometry coding category.

Which has the bigger context window?

Qwen2.5 7B Instruct does, with 131K tokens against 33K.

How many benchmarks do Qwen Max and Qwen2.5 7B Instruct share?

2 benchmarks have published results for both models. Qwen Max has 23 scored results on Noometry and Qwen2.5 7B Instruct has 15.

Related comparisons

Go deeper