Model comparison

Olmo 3.1 32b Think vs Qwen3-Coder 480B-A35B Instruct

Olmo 3.1 32b Think and Qwen3-Coder 480B-A35B Instruct score almost the same on the Noometry Index (37.9 vs 38.1), so choose on price, context window or the category you care about most.

Last verified . 15 shared benchmarks.

Summary

  • They share 15 benchmarks with published results for both. Olmo 3.1 32b Think scores higher in 1 category and Qwen3-Coder 480B-A35B Instruct in 7 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in multilingual, where Qwen3-Coder 480B-A35B Instruct leads 47.7 to 38.1.

Side by side

Olmo 3.1 32b Think and Qwen3-Coder 480B-A35B Instruct specifications
Olmo 3.1 32b ThinkQwen3-Coder 480B-A35B Instruct
ProviderAllen Institute for AI (Ai2)Alibaba (Qwen)
Noometry Index37.938.1
Released—2025-04
WeightsOpenOpen
Context window—262K
Max output—66K
Input $ / M tokens—$1.50
Output $ / M tokens—$7.50
Results tracked1525

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Olmo 3.1 32b Think leads

Olmo 3.1 32b Think: 37.7 (#189), Qwen3-Coder 480B-A35B Instruct: 35.5 (#223)

Coding benchmarks
BenchmarkOlmo 3.1 32b ThinkQwen3-Coder 480B-A35B Instruct
LMArena Coding12911412
SWE-bench Verified (bash only)—55.4%
LMArena WebDev—1275
GSO—4.9%
WeirdML—41.2%
ALE-Bench—461.45
AlgoTune—1.44

Agentic & Tool Use Not comparable

Olmo 3.1 32b Think: —, Qwen3-Coder 480B-A35B Instruct: 23.9 (#123)

Agentic & Tool Use benchmarks
BenchmarkOlmo 3.1 32b ThinkQwen3-Coder 480B-A35B Instruct
Terminal-Bench—27.2%

Reasoning Too close to call

Olmo 3.1 32b Think: 25.2 (#150), Qwen3-Coder 480B-A35B Instruct: 25.5 (#149)

Reasoning benchmarks
BenchmarkOlmo 3.1 32b ThinkQwen3-Coder 480B-A35B Instruct
LMArena Hard Prompts12721372
Kagi LLM Benchmark—49.5%

Math Qwen3-Coder 480B-A35B Instruct leads

Olmo 3.1 32b Think: 36.3 (#168), Qwen3-Coder 480B-A35B Instruct: 37.6 (#150)

Math benchmarks
BenchmarkOlmo 3.1 32b ThinkQwen3-Coder 480B-A35B Instruct
LMArena Math13051365

Knowledge Qwen3-Coder 480B-A35B Instruct leads

Olmo 3.1 32b Think: 35.7 (#181), Qwen3-Coder 480B-A35B Instruct: 37.0 (#162)

Knowledge benchmarks
BenchmarkOlmo 3.1 32b ThinkQwen3-Coder 480B-A35B Instruct
LMArena Expert12951338

Multilingual Qwen3-Coder 480B-A35B Instruct leads

Olmo 3.1 32b Think: 38.1 (#231), Qwen3-Coder 480B-A35B Instruct: 47.7 (#148)

Multilingual benchmarks
BenchmarkOlmo 3.1 32b ThinkQwen3-Coder 480B-A35B Instruct
LMArena Non-English12091346
LMArena Chinese12421357
LMArena French12601398
LMArena German12621325
LMArena Russian11931366
LMArena Spanish12891360
LMArena Japanese—1310
LMArena Korean—1305

Instruction Following Qwen3-Coder 480B-A35B Instruct leads

Olmo 3.1 32b Think: 65.6 (#218), Qwen3-Coder 480B-A35B Instruct: 71.6 (#147)

Instruction Following benchmarks
BenchmarkOlmo 3.1 32b ThinkQwen3-Coder 480B-A35B Instruct
LMArena Instruction Following12471355

Long Context Qwen3-Coder 480B-A35B Instruct leads

Olmo 3.1 32b Think: 38.6 (#195), Qwen3-Coder 480B-A35B Instruct: 42.0 (#131)

Long Context benchmarks
BenchmarkOlmo 3.1 32b ThinkQwen3-Coder 480B-A35B Instruct
LMArena Longer Query12721378

Writing & Preference Qwen3-Coder 480B-A35B Instruct leads

Olmo 3.1 32b Think: 46.2 (#220), Qwen3-Coder 480B-A35B Instruct: 55.3 (#147)

Writing & Preference benchmarks
BenchmarkOlmo 3.1 32b ThinkQwen3-Coder 480B-A35B Instruct
LMArena Text12721357
LMArena Creative Writing12261333
LMArena Multi-Turn12521365

Frequently asked questions

Is Olmo 3.1 32b Think better than Qwen3-Coder 480B-A35B Instruct?

Olmo 3.1 32b Think and Qwen3-Coder 480B-A35B Instruct score almost the same on the Noometry Index (37.9 vs 38.1), so choose on price, context window or the category you care about most.

Is Olmo 3.1 32b Think or Qwen3-Coder 480B-A35B Instruct better for coding?

Olmo 3.1 32b Think scores higher on coding benchmarks: 37.7 versus 35.5 in the Noometry coding category.

How many benchmarks do Olmo 3.1 32b Think and Qwen3-Coder 480B-A35B Instruct share?

15 benchmarks have published results for both models. Olmo 3.1 32b Think has 15 scored results on Noometry and Qwen3-Coder 480B-A35B Instruct has 25.

Related comparisons

Go deeper