Model comparison

Olmo 2 0325 32b Instruct vs Qwen Max

Qwen Max is the stronger model overall, scoring 34.7 to 32.7 on the Noometry Index.

Last verified . 11 shared benchmarks.

Summary

  • They share 11 benchmarks with published results for both. Olmo 2 0325 32b Instruct scores higher in 2 categories and Qwen Max in 6 categories; 8 gaps are clear of the uncertainty.
  • The widest gap is in knowledge, where Qwen Max leads 30.3 to 19.5.
  • Olmo 2 0325 32b Instruct has downloadable open weights; the other is API-only.

Side by side

Olmo 2 0325 32b Instruct and Qwen Max specifications
Olmo 2 0325 32b InstructQwen Max
ProviderAllen Institute for AI (Ai2)Alibaba (Qwen)
Noometry Index32.734.7
Released—2024-04-03
WeightsOpenProprietary
Context window—33K
Max output—8K
Input $ / M tokens—$1.60
Output $ / M tokens—$6.40
Results tracked1623

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Olmo 2 0325 32b Instruct leads

Olmo 2 0325 32b Instruct: 35.2 (#227), Qwen Max: 30.7 (#292)

Coding benchmarks
BenchmarkOlmo 2 0325 32b InstructQwen Max
LMArena Coding12101288
Aider Polyglot—21.8%

Reasoning Qwen Max leads

Olmo 2 0325 32b Instruct: 23.6 (#175), Qwen Max: 25.1 (#151)

Reasoning benchmarks
BenchmarkOlmo 2 0325 32b InstructQwen Max
LMArena Hard Prompts12081269

Math Olmo 2 0325 32b Instruct leads

Olmo 2 0325 32b Instruct: 26.8 (#255), Qwen Max: 22.3 (#276)

Math benchmarks
BenchmarkOlmo 2 0325 32b InstructQwen Max
LMArena Math12081275
OTIS Mock AIME 2024-2025—16.1%
Omni-MATH16.1%—
MATH Level 5—67.2%
FrontierMath (Feb 2025 set)—1%

Knowledge Qwen Max leads

Olmo 2 0325 32b Instruct: 19.5 (#279), Qwen Max: 30.3 (#228)

Knowledge benchmarks
BenchmarkOlmo 2 0325 32b InstructQwen Max
GPQA Diamond—56.1%
MMLU-Pro41.4%—
GPQA (HELM)28.7%—
LMArena Expert—1248

Multilingual Qwen Max leads

Olmo 2 0325 32b Instruct: 34.8 (#248), Qwen Max: 41.8 (#202)

Multilingual benchmarks
BenchmarkOlmo 2 0325 32b InstructQwen Max
LMArena Non-English11601263
LMArena Chinese11921254
LMArena Russian11871274
LMArena French—1330
LMArena German—1254
LMArena Japanese—1205
LMArena Korean—1142
LMArena Spanish—1290

Instruction Following Qwen Max leads

Olmo 2 0325 32b Instruct: 61.5 (#244), Qwen Max: 66.5 (#208)

Instruction Following benchmarks
BenchmarkOlmo 2 0325 32b InstructQwen Max
LMArena Instruction Following11861262
IFEval78%—

Long Context Qwen Max leads

Olmo 2 0325 32b Instruct: 36.2 (#234), Qwen Max: 39.4 (#180)

Long Context benchmarks
BenchmarkOlmo 2 0325 32b InstructQwen Max
LMArena Longer Query11941288
Fiction.LiveBench—66.7%

Writing & Preference Qwen Max leads

Olmo 2 0325 32b Instruct: 42.1 (#236), Qwen Max: 47.8 (#205)

Writing & Preference benchmarks
BenchmarkOlmo 2 0325 32b InstructQwen Max
LMArena Text12181282
LMArena Creative Writing11991248
LMArena Multi-Turn12211277
WildBench73.4%—

Frequently asked questions

Is Olmo 2 0325 32b Instruct better than Qwen Max?

Qwen Max is the stronger model overall, scoring 34.7 to 32.7 on the Noometry Index.

Is Olmo 2 0325 32b Instruct or Qwen Max better for coding?

Olmo 2 0325 32b Instruct scores higher on coding benchmarks: 35.2 versus 30.7 in the Noometry coding category.

How many benchmarks do Olmo 2 0325 32b Instruct and Qwen Max share?

11 benchmarks have published results for both models. Olmo 2 0325 32b Instruct has 16 scored results on Noometry and Qwen Max has 23.

Related comparisons

Go deeper