Alibaba (Qwen), open weights
Qwen3 32B
Qwen3 32B by Alibaba (Qwen) ranks 172nd of 354 ranked models on the Noometry Index as of October 2026, with a score of 39.2. Its strongest category is agentic & tool use, where it ranks 62nd. API pricing starts at $0.70 per million input tokens and $2.80 per million output tokens, with a 131K-token context window.
Last verified
Specifications
- Noometry rank
- #172 of 354
- Index score
- 39.2
- Evidence
- Confirmed 26 results
- Provider
Alibaba (Qwen)
- Released
- April 1, 2025
- Weights
- Open weights
- Reasoning
- Yes
- Context window
- 131K
- Max output
- 16K
- Input price
- $0.70 / M
- Output price
- $2.80 / M
- Blended price
- $1.22 / M
- Output speed
- 86 tokens/s Kagi
- Value
- #130 of 219
- Knowledge cutoff
- April 2025
- Input
- text
- Hugging Face
- Qwen/Qwen3-32B
Category scores
Each category score combines every public result we have in that category.
- Coding 37.7
- Agentic & Tool Use 32.6
- Reasoning 20.2
- Math 39.7
- Knowledge 40.0
- Multilingual 45.6
- Instruction Following 68.9
- Long Context 43.8
- Writing & Preference 52.9
| Category | Score | Rank | Results |
|---|---|---|---|
| Coding | 37.7 | #190 | 3 |
| Agentic & Tool Use | 32.6 | #62 | 1 |
| Reasoning | 20.2 | #241 | 6 |
| Math | 39.7 | #99 | 2 |
| Knowledge | 40.0 | #125 | 3 |
| Multilingual | 45.6 | #167 | 1 |
| Instruction Following | 68.9 | #179 | 1 |
| Long Context | 43.8 | #87 | 2 |
| Writing & Preference | 52.9 | #163 | 3 |
Strengths and weaknesses
Categories where Qwen3 32B places highest and lowest among the models ranked in each, with its score against that category's median.
Strongest categories
| Category | Score | vs median | Rank |
|---|---|---|---|
| Long Context | 43.8 | +2.8 | #87 of 296, top 30% |
| Math | 39.7 | +3.1 | #99 of 327, top 31% |
| Knowledge | 40.0 | +2.7 | #125 of 314, top 40% |
Weakest categories
| Category | Score | vs median | Rank |
|---|---|---|---|
| Reasoning | 20.2 | −3.4 | #241 of 350, top 69% |
| Instruction Following | 68.9 | −2.4 | #179 of 305, top 59% |
| Multilingual | 45.6 | −1.8 | #167 of 297, top 57% |
Closest competitors
The models ranked just above and below Qwen3 32B. When scores are this close, price and speed are often the better way to choose.
| Model | Rank | Score | Blended $/M | Speed | |
|---|---|---|---|---|---|
| Olmo 3.1 32b Instruct | #168 | 39.4 | — | — | Compare |
| Granite 4.2 3b | #169 | 39.4 | — | — | Compare |
| Gemini 2.5 Flash | #170 | 39.3 | $0.85 | 152 | Compare |
| Step 2 16k Exp 202412 | #171 | 39.2 | — | — | Compare |
| Gemini 2.0 Pro | #173 | 39.1 | — | — | Compare |
| Molmo 2 8b | #174 | 39.1 | — | — | Compare |
| Mercury 2 | #175 | 39.1 | $0.38 | — | Compare |
| Mistral Large 3 | #176 | 39.1 | $0.38 | 7 | Compare |
Sponsored placements are available on pages like this one. Advertise on Noometry
Benchmark results
Every published result we track, with its source. Bold rows are the ones used for ranking; where several exist we prefer independent runs over self-reported numbers.
Coding
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| Aider Polyglot | 40% | #28 of 44, top 64% | Epoch AI | ||
| SciCode | 35.4% | #99 of 121, top 82% | Epoch AI | ||
| LMArena Coding | 1358 | #170 of 294, top 58% | LMArena | 2026-10-08 |
Agentic & Tool Use
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| Berkeley Function Calling Leaderboard | 48.7% | #20 of 49, top 41% | fc | Berkeley Function Calling Leaderboard |
Reasoning
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| Kagi LLM Benchmark | 54.9% | #53 of 99, top 54% | Kagi LLM Benchmark | ||
| Kagi LLM Benchmark | 48.7% | Kagi LLM Benchmark | |||
| CritPt | 0.3% | #98 of 134, top 74% | Epoch AI | ||
| Chess Puzzles | 5% | #94 of 129, top 73% | Epoch AI | 2026-08-28 | |
| Chess Puzzles | 1% | none | Epoch AI | 2026-08-28 | |
| LMArena Hard Prompts | 1334 | #170 of 297, top 58% | LMArena | 2026-10-08 | |
| DTBench | 67.5% | #99 of 151, top 66% | Epoch AI | ||
| LMCA | 17.3% | #98 of 125, top 79% | Epoch AI | ||
| Epoch Capabilities Index | 138.51 | #113 of 213, top 54% | Epoch AI | 2025-04-29 |
Math
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| OTIS Mock AIME 2024-2025 | 66.9% | #96 of 173, top 56% | Epoch AI | 2026-08-30 | |
| OTIS Mock AIME 2024-2025 | 23.1% | none | Epoch AI | 2026-08-30 | |
| LMArena Math | 1399 | #126 of 285, top 45% | LMArena | 2026-10-08 |
Knowledge
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| GPQA Diamond | 65.7% | #105 of 186, top 57% | Epoch AI | 2026-08-28 | |
| GPQA Diamond | 54.1% | none | Epoch AI | 2026-08-30 | |
| Vectara Hallucination Rate (lower is better) | 5.9% | #21 of 96, top 22% | Vectara Hallucination Leaderboard | ||
| LMArena Expert | 1362 | #146 of 273, top 54% | LMArena | 2026-10-08 |
Multilingual
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Non-English | 1317 | #167 of 297, top 57% | LMArena | 2026-10-08 | |
| LMArena Chinese | 1357 | #163 of 285, top 58% | LMArena | 2026-10-08 | |
| LMArena German | 1341 | #136 of 231, top 59% | LMArena | 2026-10-08 | |
| LMArena Russian | 1311 | #171 of 283, top 61% | LMArena | 2026-10-08 |
Instruction Following
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Instruction Following | 1305 | #173 of 298, top 59% | LMArena | 2026-10-08 |
Long Context
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| Fiction.LiveBench | 74.2% | #16 of 47, top 35% | Epoch AI | ||
| LMArena Longer Query | 1327 | #165 of 291, top 57% | LMArena | 2026-10-08 |
Writing & Preference
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Text | 1340 | #163 of 297, top 55% | LMArena | 2026-10-08 | |
| LMArena Creative Writing | 1297 | #167 of 295, top 57% | LMArena | 2026-10-08 | |
| LMArena Multi-Turn | 1331 | #172 of 295, top 59% | LMArena | 2026-10-08 |
API pricing by provider
| Route | Input $/M | Output $/M | Cached input $/M | Checked |
|---|---|---|---|---|
| alibaba | $0.70 | $2.80 | — | 2026-10-10 |
| bedrock | $0.15 | $0.60 | — | 2026-10-10 |
| deepinfra | $0.08 | $0.28 | — | 2026-10-10 |
| openrouter | $0.08 | $0.28 | — | 2026-10-10 |
Compare Qwen3 32B
- Qwen3 32B vs Qwen2-72B
- Qwen3 32B vs Step 2 16k Exp 202412
- Qwen3 32B vs Gemini 2.0 Pro
- Qwen3 32B vs Gemini 2.5 Flash
- Qwen3 32B vs Molmo 2 8b
- Qwen3 32B vs Granite 4.2 3b
- Qwen3 32B vs Mercury 2
- Qwen3 32B vs GPT-6 Astra
- Qwen3 32B vs Claude Fable 5.1
- Qwen3 32B vs Gemini 3.8 Flash
- Qwen3 32B vs Kimi K3
- Qwen3 32B vs Grok 4.6
- Qwen3 32B vs GLM-5.3
- Qwen3 32B vs Muse Spark 1.3
Other Alibaba (Qwen) models
- Qwen3.8 Max56.8
- Qwen3.7 Max51.5
- Qwen3.6 Max Preview51.5
- Qwen3.6 Plus47.5
- Qwen3.5 397B-A17B46.0
- Qwen3.8 27B46.0
- Qwen3.5 Max Preview45.3
- Qwen3.7 Plus45.3
Frequently asked questions
How good is Qwen3 32B?
Qwen3 32B by Alibaba (Qwen) ranks 172nd of 354 ranked models on the Noometry Index as of October 2026, with a score of 39.2. Its strongest category is agentic & tool use, where it ranks 62nd. API pricing starts at $0.70 per million input tokens and $2.80 per million output tokens, with a 131K-token context window.
How much does Qwen3 32B cost?
Qwen3 32B costs $0.70 per million input tokens and $2.80 per million output tokens on Alibaba (Qwen)'s own API.
What is Qwen3 32B's context window?
Qwen3 32B accepts up to 131K tokens of input and can write up to 16K tokens in one response.
Is Qwen3 32B open source?
Yes. Qwen3 32B's weights are downloadable from Hugging Face (Qwen/Qwen3-32B); check the license for commercial terms.
How fast is Qwen3 32B?
Qwen3 32B generated about 86 output tokens per second in the Kagi LLM Benchmark's timed runs. Speed varies by provider, load and reasoning effort.
What are Qwen3 32B's strengths and weaknesses?
Relative to other ranked models, Qwen3 32B places best in long context, math, knowledge and lowest in reasoning, instruction following, multilingual.
What is Qwen3 32B best at?
Its best category is agentic & tool use, where it ranks 62nd on Noometry.