Model comparison

Olmo 3.1 32b Instruct vs Qwen3 32B

Olmo 3.1 32b Instruct and Qwen3 32B score almost the same on the Noometry Index (39.4 vs 39.2), so choose on price, context window or the category you care about most.

Last verified . 13 shared benchmarks.

Qwen3 32B Alibaba (Qwen)

39.2

Rank #172 Confirmed

Summary

  • They share 13 benchmarks with published results for both. Olmo 3.1 32b Instruct scores higher in 2 categories and Qwen3 32B in 6 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where Olmo 3.1 32b Instruct leads 26.4 to 20.2.

Side by side

Olmo 3.1 32b Instruct and Qwen3 32B specifications
Olmo 3.1 32b InstructQwen3 32B
ProviderAllen Institute for AI (Ai2)Alibaba (Qwen)
Noometry Index39.439.2
Released—2025-04
WeightsOpenOpen
Context window—131K
Max output—16K
Input $ / M tokens—$0.70
Output $ / M tokens—$2.80
Results tracked1626

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Olmo 3.1 32b Instruct leads

Olmo 3.1 32b Instruct: 39.5 (#157), Qwen3 32B: 37.7 (#190)

Coding benchmarks
BenchmarkOlmo 3.1 32b InstructQwen3 32B
LMArena Coding13471358
Aider Polyglot—40%
SciCode—35.4%

Agentic & Tool Use Not comparable

Olmo 3.1 32b Instruct: —, Qwen3 32B: 32.6 (#62)

Agentic & Tool Use benchmarks
BenchmarkOlmo 3.1 32b InstructQwen3 32B
Berkeley Function Calling Leaderboard—48.7%

Reasoning Olmo 3.1 32b Instruct leads

Olmo 3.1 32b Instruct: 26.4 (#132), Qwen3 32B: 20.2 (#241)

Reasoning benchmarks
BenchmarkOlmo 3.1 32b InstructQwen3 32B
LMArena Hard Prompts13221334
Kagi LLM Benchmark—54.9%
CritPt—0.3%
Chess Puzzles—5%
DTBench—67.5%
LMCA—17.3%
Epoch Capabilities Index—138.51

Math Qwen3 32B leads

Olmo 3.1 32b Instruct: 36.3 (#167), Qwen3 32B: 39.7 (#99)

Math benchmarks
BenchmarkOlmo 3.1 32b InstructQwen3 32B
LMArena Math13051399
OTIS Mock AIME 2024-2025—66.9%

Knowledge Qwen3 32B leads

Olmo 3.1 32b Instruct: 36.1 (#175), Qwen3 32B: 40.0 (#125)

Knowledge benchmarks
BenchmarkOlmo 3.1 32b InstructQwen3 32B
LMArena Expert13081362
GPQA Diamond—65.7%
Vectara Hallucination Rate—5.9%

Multilingual Qwen3 32B leads

Olmo 3.1 32b Instruct: 42.6 (#191), Qwen3 32B: 45.6 (#167)

Multilingual benchmarks
BenchmarkOlmo 3.1 32b InstructQwen3 32B
LMArena Non-English12751317
LMArena Chinese13041357
LMArena German12821341
LMArena Russian12681311
LMArena French1328—
LMArena Korean1206—
LMArena Spanish1336—

Instruction Following Too close to call

Olmo 3.1 32b Instruct: 68.6 (#187), Qwen3 32B: 68.9 (#179)

Instruction Following benchmarks
BenchmarkOlmo 3.1 32b InstructQwen3 32B
LMArena Instruction Following12991305

Long Context Qwen3 32B leads

Olmo 3.1 32b Instruct: 39.9 (#166), Qwen3 32B: 43.8 (#87)

Long Context benchmarks
BenchmarkOlmo 3.1 32b InstructQwen3 32B
LMArena Longer Query13121327
Fiction.LiveBench—74.2%

Writing & Preference Qwen3 32B leads

Olmo 3.1 32b Instruct: 50.2 (#185), Qwen3 32B: 52.9 (#163)

Writing & Preference benchmarks
BenchmarkOlmo 3.1 32b InstructQwen3 32B
LMArena Text13111340
LMArena Creative Writing12641297
LMArena Multi-Turn13091331

Frequently asked questions

Is Olmo 3.1 32b Instruct better than Qwen3 32B?

Olmo 3.1 32b Instruct and Qwen3 32B score almost the same on the Noometry Index (39.4 vs 39.2), so choose on price, context window or the category you care about most.

Is Olmo 3.1 32b Instruct or Qwen3 32B better for coding?

Olmo 3.1 32b Instruct scores higher on coding benchmarks: 39.5 versus 37.7 in the Noometry coding category.

How many benchmarks do Olmo 3.1 32b Instruct and Qwen3 32B share?

13 benchmarks have published results for both models. Olmo 3.1 32b Instruct has 16 scored results on Noometry and Qwen3 32B has 26.

Related comparisons

Go deeper