Alibaba (Qwen), open weights
Qwen3 235B-A22B
Qwen3 235B-A22B by Alibaba (Qwen) ranks 91st of 354 ranked models on the Noometry Index as of October 2026, with a score of 43.5. Its strongest category is long context, where it ranks 26th. API pricing starts at $0.70 per million input tokens and $2.80 per million output tokens, with a 131K-token context window.
Last verified
Specifications
- Noometry rank
- #91 of 354
- Index score
- 43.5
- Evidence
- Confirmed 49 results
- Provider
Alibaba (Qwen)
- Released
- April 1, 2025
- Weights
- Open weights
- Reasoning
- Yes
- Context window
- 131K
- Max output
- 16K
- Input price
- $0.70 / M
- Output price
- $2.80 / M
- Blended price
- $1.22 / M
- Output speed
- 85 tokens/s Kagi
- Value
- #126 of 219
- Knowledge cutoff
- April 2025
- Input
- text
- Hugging Face
- Qwen/Qwen3-235B-A22B-Instruct-2507
Category scores
Each category score combines every public result we have in that category.
- Coding 44.3
- Agentic & Tool Use 33.9
- Reasoning 15.7
- Math 50.4
- Knowledge 49.6
- Multilingual 52.3
- Instruction Following 72.6
- Long Context 46.1
- Writing & Preference 59.6
| Category | Score | Rank | Results |
|---|---|---|---|
| Coding | 44.3 | #75 | 4 |
| Agentic & Tool Use | 33.9 | #51 | 1 |
| Reasoning | 15.7 | #311 | 10 |
| Math | 50.4 | #57 | 4 |
| Knowledge | 49.6 | #73 | 7 |
| Multilingual | 52.3 | #89 | 1 |
| Instruction Following | 72.6 | #136 | 2 |
| Long Context | 46.1 | #26 | 2 |
| Writing & Preference | 59.6 | #108 | 6 |
Strengths and weaknesses
Categories where Qwen3 235B-A22B places highest and lowest among the models ranked in each, with its score against that category's median.
Strongest categories
| Category | Score | vs median | Rank |
|---|---|---|---|
| Long Context | 46.1 | +5.2 | #26 of 296, top 9% |
| Math | 50.4 | +13.8 | #57 of 327, top 18% |
| Coding | 44.3 | +5.5 | #75 of 340, top 23% |
Weakest categories
| Category | Score | vs median | Rank |
|---|---|---|---|
| Reasoning | 15.7 | −7.9 | #311 of 350, top 89% |
| Instruction Following | 72.6 | +1.3 | #136 of 305, top 45% |
| Writing & Preference | 59.6 | +5.8 | #108 of 312, top 35% |
Closest competitors
The models ranked just above and below Qwen3 235B-A22B. When scores are this close, price and speed are often the better way to choose.
| Model | Rank | Score | Blended $/M | Speed | |
|---|---|---|---|---|---|
| Qwen3 Max | #87 | 43.7 | $2.40 | 48 | Compare |
| MiMo-V2-Omni | #88 | 43.6 | $0.18 | — | Compare |
| Kimi K2.5 Instant | #89 | 43.6 | — | — | Compare |
| Gemma 4 31B IT | #90 | 43.5 | $0.15 | 3 | Compare |
| Gemma 4 26B A4B IT | #92 | 43.5 | $0.11 | — | Compare |
| MiMo-V2.5 | #93 | 43.4 | $0.18 | — | Compare |
| Kimi K2.7 Code | #94 | 43.3 | $1.71 | — | Compare |
| Qwen3-VL 235B-A22B | #95 | 43.2 | $1.22 | — | Compare |
Sponsored placements are available on pages like this one. Advertise on Noometry
Benchmark results
Every published result we track, with its source. Bold rows are the ones used for ranking; where several exist we prefer independent runs over self-reported numbers.
Coding
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| Aider Polyglot | 59.6% | #14 of 44, top 32% | Epoch AI | ||
| Aider Polyglot | 59.6% | #14 of 44, top 32% | Epoch AI | ||
| SciCode | 42.4% | #70 of 121, top 58% | Epoch AI | ||
| WeirdML | 38.7% | Epoch AI | |||
| WeirdML | 41% | #74 of 119, top 63% | Epoch AI | ||
| WeirdML | 38.7% | Epoch AI | |||
| WeirdML | 37.3% | Epoch AI | |||
| LMArena Coding | 1445 | #94 of 294, top 32% | LMArena | 2026-10-08 | |
| LMArena Coding | 1424 | LMArena | 2026-10-08 | ||
| LMArena Coding | 1398 | no-thinking | LMArena | 2026-10-08 |
Agentic & Tool Use
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| Berkeley Function Calling Leaderboard | 52.1% | #17 of 49, top 35% | prompt | Berkeley Function Calling Leaderboard | |
| Vending-Bench 2 | -11.34 | #57 of 60, top 95% | Epoch AI |
Reasoning
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| ARC-AGI-2 | 1.3% | #68 of 83, top 82% | Epoch AI | ||
| SimpleBench | 31% | #60 of 77, top 78% | Epoch AI | ||
| Kagi LLM Benchmark | 55% | Kagi LLM Benchmark | |||
| Kagi LLM Benchmark | 69.4% | #26 of 99, top 27% | Kagi LLM Benchmark | ||
| ARC-AGI-1 | 11% | #73 of 83, top 88% | Epoch AI | ||
| CritPt | 0% | #131 of 134, top 98% | Epoch AI | ||
| Chess Puzzles | 12% | #80 of 129, top 63% | Epoch AI | 2025-12-11 | |
| LMArena Hard Prompts | 1433 | #88 of 297, top 30% | LMArena | 2026-10-08 | |
| LMArena Hard Prompts | 1416 | LMArena | 2026-10-08 | ||
| LMArena Hard Prompts | 1393 | no-thinking | LMArena | 2026-10-08 | |
| Mystery Game Puzzles | 9% | #63 of 74, top 86% | Epoch AI | 2026-08-27 | |
| DTBench | 78.4% | Epoch AI | |||
| DTBench | 75.7% | Epoch AI | |||
| DTBench | 80.3% | #74 of 151, top 50% | Epoch AI | ||
| LMCA | 25% | Epoch AI | |||
| LMCA | 29.3% | #76 of 125, top 61% | Epoch AI | ||
| LMCA | 23.5% | Epoch AI | |||
| Epoch Capabilities Index | 143.85 | #90 of 213, top 43% | Epoch AI | 2025-07-25 | |
| Epoch Capabilities Index | 138.92 | Epoch AI | 2025-07-25 | ||
| Epoch Capabilities Index | 139.35 | Epoch AI | 2025-04-28 | ||
| ForecastBench | 59.7 | #36 of 72, top 50% | Epoch AI |
Math
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| OTIS Mock AIME 2024-2025 | 86.7% | #63 of 173, top 37% | Epoch AI | 2025-12-10 | |
| Omni-MATH | 71.8% | #3 of 57, top 6% | HELM Capabilities | ||
| Omni-MATH | 54.8% | HELM Capabilities | |||
| LMArena Math | 1432 | #81 of 285, top 29% | LMArena | 2026-10-08 | |
| LMArena Math | 1412 | LMArena | 2026-10-08 | ||
| LMArena Math | 1397 | no-thinking | LMArena | 2026-10-08 | |
| MATH Level 5 | 68.9% | #32 of 79, top 41% | Epoch AI | 2025-06-03 | |
| FrontierMath (Feb 2025 set) | 8.5% | #39 of 68, top 58% | Epoch AI | 2025-12-11 | |
| FrontierMath Tier 4 (v1) | 0% | #53 of 55, top 97% | Epoch AI | 2025-12-11 |
Knowledge
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| GPQA Diamond | 80.1% | #79 of 186, top 43% | Epoch AI | 2025-12-11 | |
| GPQA Diamond | 70.7% | Epoch AI | 2025-06-03 | ||
| SimpleQA Verified | 40.4% | #43 of 77, top 56% | Epoch AI | 2026-08-27 | |
| MMLU-Pro | 81.7% | HELM Capabilities | |||
| MMLU-Pro | 84.4% | #8 of 58, top 14% | HELM Capabilities | ||
| Confabulations (lower is better) | 16.8% | Lech Mazur benchmarks | |||
| Confabulations (lower is better) | 15.6% | #20 of 51, top 40% | Lech Mazur benchmarks | ||
| Vectara Hallucination Rate (lower is better) | 9.3% | #47 of 96, top 49% | Vectara Hallucination Leaderboard | ||
| GPQA (HELM) | 72.7% | #8 of 57, top 15% | HELM Capabilities | ||
| GPQA (HELM) | 62.3% | HELM Capabilities | |||
| LMArena Expert | 1442 | LMArena | 2026-10-08 | ||
| LMArena Expert | 1463 | #58 of 273, top 22% | LMArena | 2026-10-08 | |
| LMArena Expert | 1369 | no-thinking | LMArena | 2026-10-08 |
Multilingual
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Non-English | 1409 | #88 of 297, top 30% | LMArena | 2026-10-08 | |
| LMArena Non-English | 1397 | LMArena | 2026-10-08 | ||
| LMArena Non-English | 1384 | no-thinking | LMArena | 2026-10-08 | |
| LMArena Chinese | 1465 | LMArena | 2026-10-08 | ||
| LMArena Chinese | 1481 | #62 of 285, top 22% | LMArena | 2026-10-08 | |
| LMArena Chinese | 1425 | no-thinking | LMArena | 2026-10-08 | |
| LMArena French | 1445 | #80 of 223, top 36% | LMArena | 2026-10-08 | |
| LMArena French | 1369 | no-thinking | LMArena | 2026-10-08 | |
| LMArena German | 1387 | LMArena | 2026-10-08 | ||
| LMArena German | 1408 | LMArena | 2026-10-08 | ||
| LMArena German | 1433 | #62 of 231, top 27% | LMArena | 2026-10-08 | |
| LMArena Japanese | 1386 | LMArena | 2026-10-08 | ||
| LMArena Japanese | 1399 | #60 of 211, top 29% | LMArena | 2026-10-08 | |
| LMArena Japanese | 1371 | no-thinking | LMArena | 2026-10-08 | |
| LMArena Korean | 1391 | #68 of 213, top 32% | LMArena | 2026-10-08 | |
| LMArena Korean | 1375 | LMArena | 2026-10-08 | ||
| LMArena Korean | 1358 | no-thinking | LMArena | 2026-10-08 | |
| LMArena Russian | 1400 | LMArena | 2026-10-08 | ||
| LMArena Russian | 1411 | #94 of 283, top 34% | LMArena | 2026-10-08 | |
| LMArena Russian | 1394 | no-thinking | LMArena | 2026-10-08 | |
| LMArena Spanish | 1389 | LMArena | 2026-10-08 | ||
| LMArena Spanish | 1430 | #83 of 226, top 37% | LMArena | 2026-10-08 | |
| LMArena Spanish | 1403 | no-thinking | LMArena | 2026-10-08 |
Instruction Following
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| IFEval | 81.6% | HELM Capabilities | |||
| IFEval | 83.5% | #28 of 57, top 50% | HELM Capabilities | ||
| LMArena Instruction Following | 1386 | LMArena | 2026-10-08 | ||
| LMArena Instruction Following | 1408 | #90 of 298, top 31% | LMArena | 2026-10-08 | |
| LMArena Instruction Following | 1362 | no-thinking | LMArena | 2026-10-08 |
Long Context
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| Fiction.LiveBench | 75% | #15 of 47, top 32% | Epoch AI | ||
| Fiction.LiveBench | 67.7% | Epoch AI | |||
| Fiction.LiveBench | 52.9% | Epoch AI | |||
| LMArena Longer Query | 1400 | LMArena | 2026-10-08 | ||
| LMArena Longer Query | 1426 | #84 of 291, top 29% | LMArena | 2026-10-08 | |
| LMArena Longer Query | 1389 | no-thinking | LMArena | 2026-10-08 |
Writing & Preference
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Text | 1415 | LMArena | 2026-10-08 | ||
| LMArena Text | 1419 | #96 of 297, top 33% | LMArena | 2026-10-08 | |
| LMArena Text | 1394 | no-thinking | LMArena | 2026-10-08 | |
| LMArena Creative Writing | 1384 | #102 of 295, top 35% | LMArena | 2026-10-08 | |
| LMArena Creative Writing | 1375 | LMArena | 2026-10-08 | ||
| LMArena Creative Writing | 1356 | no-thinking | LMArena | 2026-10-08 | |
| Short-Story Creative Writing | 82.4% | Epoch AI | |||
| Short-Story Creative Writing | 83% | #10 of 39, top 26% | Epoch AI | ||
| EQ-Bench Creative Writing | 1366 | #73 of 115, top 64% | EQ-Bench | ||
| WildBench | 86.6% | Best of 57 | HELM Capabilities | ||
| WildBench | 82.8% | HELM Capabilities | |||
| LMArena Multi-Turn | 1432 | #81 of 295, top 28% | LMArena | 2026-10-08 | |
| LMArena Multi-Turn | 1406 | LMArena | 2026-10-08 | ||
| LMArena Multi-Turn | 1398 | no-thinking | LMArena | 2026-10-08 |
API pricing by provider
Compare Qwen3 235B-A22B
- Qwen3 235B-A22B vs Qwen2-72B
- Qwen3 235B-A22B vs Gemma 4 31B IT
- Qwen3 235B-A22B vs Gemma 4 26B A4B IT
- Qwen3 235B-A22B vs Kimi K2.5 Instant
- Qwen3 235B-A22B vs MiMo-V2.5
- Qwen3 235B-A22B vs MiMo-V2-Omni
- Qwen3 235B-A22B vs Kimi K2.7 Code
- Qwen3 235B-A22B vs GPT-6 Astra
- Qwen3 235B-A22B vs Claude Fable 5.1
- Qwen3 235B-A22B vs Gemini 3.8 Flash
- Qwen3 235B-A22B vs Kimi K3
- Qwen3 235B-A22B vs Grok 4.6
- Qwen3 235B-A22B vs GLM-5.3
- Qwen3 235B-A22B vs Muse Spark 1.3
Other Alibaba (Qwen) models
- Qwen3.8 Max56.8
- Qwen3.7 Max51.5
- Qwen3.6 Max Preview51.5
- Qwen3.6 Plus47.5
- Qwen3.5 397B-A17B46.0
- Qwen3.8 27B46.0
- Qwen3.5 Max Preview45.3
- Qwen3.7 Plus45.3
Frequently asked questions
How good is Qwen3 235B-A22B?
Qwen3 235B-A22B by Alibaba (Qwen) ranks 91st of 354 ranked models on the Noometry Index as of October 2026, with a score of 43.5. Its strongest category is long context, where it ranks 26th. API pricing starts at $0.70 per million input tokens and $2.80 per million output tokens, with a 131K-token context window.
How much does Qwen3 235B-A22B cost?
Qwen3 235B-A22B costs $0.70 per million input tokens and $2.80 per million output tokens on Alibaba (Qwen)'s own API.
What is Qwen3 235B-A22B's context window?
Qwen3 235B-A22B accepts up to 131K tokens of input and can write up to 16K tokens in one response.
Is Qwen3 235B-A22B open source?
Yes. Qwen3 235B-A22B's weights are downloadable from Hugging Face (Qwen/Qwen3-235B-A22B-Instruct-2507); check the license for commercial terms.
How fast is Qwen3 235B-A22B?
Qwen3 235B-A22B generated about 85 output tokens per second in the Kagi LLM Benchmark's timed runs. Speed varies by provider, load and reasoning effort.
What are Qwen3 235B-A22B's strengths and weaknesses?
Relative to other ranked models, Qwen3 235B-A22B places best in long context, math, coding and lowest in reasoning, instruction following, writing & preference.
What is Qwen3 235B-A22B best at?
Its best category is long context, where it ranks 26th on Noometry.