Alibaba (Qwen), open weights

Qwen3 235B-A22B

Qwen3 235B-A22B by Alibaba (Qwen) ranks 91st of 354 ranked models on the Noometry Index as of October 2026, with a score of 43.5. Its strongest category is long context, where it ranks 26th. API pricing starts at $0.70 per million input tokens and $2.80 per million output tokens, with a 131K-token context window.

Last verified

Specifications

Noometry rank
#91 of 354
Index score
43.5
Evidence
Confirmed 49 results
Released
April 1, 2025
Weights
Open weights
Reasoning
Yes
Context window
131K
Max output
16K
Input price
$0.70 / M
Output price
$2.80 / M
Blended price
$1.22 / M
Output speed
85 tokens/s Kagi
Value
#126 of 219
Knowledge cutoff
April 2025
Input
text

Category scores

Each category score combines every public result we have in that category.

Qwen3 235B-A22B category scores
  1. Coding 44.3
  2. Agentic & Tool Use 33.9
  3. Reasoning 15.7
  4. Math 50.4
  5. Knowledge 49.6
  6. Multilingual 52.3
  7. Instruction Following 72.6
  8. Long Context 46.1
  9. Writing & Preference 59.6
Qwen3 235B-A22B category ranks
CategoryScoreRankResults
Coding44.3#754
Agentic & Tool Use33.9#511
Reasoning15.7#31110
Math50.4#574
Knowledge49.6#737
Multilingual52.3#891
Instruction Following72.6#1362
Long Context46.1#262
Writing & Preference59.6#1086

Strengths and weaknesses

Categories where Qwen3 235B-A22B places highest and lowest among the models ranked in each, with its score against that category's median.

Strongest categories

Qwen3 235B-A22B: strongest categories
CategoryScorevs medianRank
Long Context46.1+5.2#26 of 296, top 9%
Math50.4+13.8#57 of 327, top 18%
Coding44.3+5.5#75 of 340, top 23%

Weakest categories

Qwen3 235B-A22B: weakest categories
CategoryScorevs medianRank
Reasoning15.7−7.9#311 of 350, top 89%
Instruction Following72.6+1.3#136 of 305, top 45%
Writing & Preference59.6+5.8#108 of 312, top 35%

Closest competitors

The models ranked just above and below Qwen3 235B-A22B. When scores are this close, price and speed are often the better way to choose.

Models ranked closest to Qwen3 235B-A22B
ModelRankScoreBlended $/MSpeed
Qwen3 Max#8743.7$2.4048Compare
MiMo-V2-Omni#8843.6$0.18—Compare
Kimi K2.5 Instant#8943.6——Compare
Gemma 4 31B IT#9043.5$0.153Compare
Gemma 4 26B A4B IT#9243.5$0.11—Compare
MiMo-V2.5#9343.4$0.18—Compare
Kimi K2.7 Code#9443.3$1.71—Compare
Qwen3-VL 235B-A22B#9543.2$1.22—Compare

Sponsored placements are available on pages like this one. Advertise on Noometry

Benchmark results

Every published result we track, with its source. Bold rows are the ones used for ranking; where several exist we prefer independent runs over self-reported numbers.

Coding

Qwen3 235B-A22B Coding benchmark results
BenchmarkScorePositionSettingSourceDate
Aider Polyglot59.6%#14 of 44, top 32%Epoch AI
Aider Polyglot59.6%#14 of 44, top 32%Epoch AI
SciCode42.4%#70 of 121, top 58%Epoch AI
WeirdML38.7%Epoch AI
WeirdML41%#74 of 119, top 63%Epoch AI
WeirdML38.7%Epoch AI
WeirdML37.3%Epoch AI
LMArena Coding1445#94 of 294, top 32%LMArena2026-10-08
LMArena Coding1424LMArena2026-10-08
LMArena Coding1398no-thinkingLMArena2026-10-08

Agentic & Tool Use

Qwen3 235B-A22B Agentic & Tool Use benchmark results
BenchmarkScorePositionSettingSourceDate
Berkeley Function Calling Leaderboard52.1%#17 of 49, top 35%promptBerkeley Function Calling Leaderboard
Vending-Bench 2-11.34#57 of 60, top 95%Epoch AI

Reasoning

Qwen3 235B-A22B Reasoning benchmark results
BenchmarkScorePositionSettingSourceDate
ARC-AGI-21.3%#68 of 83, top 82%Epoch AI
SimpleBench31%#60 of 77, top 78%Epoch AI
Kagi LLM Benchmark55%Kagi LLM Benchmark
Kagi LLM Benchmark69.4%#26 of 99, top 27%Kagi LLM Benchmark
ARC-AGI-111%#73 of 83, top 88%Epoch AI
CritPt0%#131 of 134, top 98%Epoch AI
Chess Puzzles12%#80 of 129, top 63%Epoch AI2025-12-11
LMArena Hard Prompts1433#88 of 297, top 30%LMArena2026-10-08
LMArena Hard Prompts1416LMArena2026-10-08
LMArena Hard Prompts1393no-thinkingLMArena2026-10-08
Mystery Game Puzzles9%#63 of 74, top 86%Epoch AI2026-08-27
DTBench78.4%Epoch AI
DTBench75.7%Epoch AI
DTBench80.3%#74 of 151, top 50%Epoch AI
LMCA25%Epoch AI
LMCA29.3%#76 of 125, top 61%Epoch AI
LMCA23.5%Epoch AI
Epoch Capabilities Index143.85#90 of 213, top 43%Epoch AI2025-07-25
Epoch Capabilities Index138.92Epoch AI2025-07-25
Epoch Capabilities Index139.35Epoch AI2025-04-28
ForecastBench59.7#36 of 72, top 50%Epoch AI

Math

Qwen3 235B-A22B Math benchmark results
BenchmarkScorePositionSettingSourceDate
OTIS Mock AIME 2024-202586.7%#63 of 173, top 37%Epoch AI2025-12-10
Omni-MATH71.8%#3 of 57, top 6%HELM Capabilities
Omni-MATH54.8%HELM Capabilities
LMArena Math1432#81 of 285, top 29%LMArena2026-10-08
LMArena Math1412LMArena2026-10-08
LMArena Math1397no-thinkingLMArena2026-10-08
MATH Level 568.9%#32 of 79, top 41%Epoch AI2025-06-03
FrontierMath (Feb 2025 set)8.5%#39 of 68, top 58%Epoch AI2025-12-11
FrontierMath Tier 4 (v1)0%#53 of 55, top 97%Epoch AI2025-12-11

Knowledge

Qwen3 235B-A22B Knowledge benchmark results
BenchmarkScorePositionSettingSourceDate
GPQA Diamond80.1%#79 of 186, top 43%Epoch AI2025-12-11
GPQA Diamond70.7%Epoch AI2025-06-03
SimpleQA Verified40.4%#43 of 77, top 56%Epoch AI2026-08-27
MMLU-Pro81.7%HELM Capabilities
MMLU-Pro84.4%#8 of 58, top 14%HELM Capabilities
Confabulations (lower is better)16.8%Lech Mazur benchmarks
Confabulations (lower is better)15.6%#20 of 51, top 40%Lech Mazur benchmarks
Vectara Hallucination Rate (lower is better)9.3%#47 of 96, top 49%Vectara Hallucination Leaderboard
GPQA (HELM)72.7%#8 of 57, top 15%HELM Capabilities
GPQA (HELM)62.3%HELM Capabilities
LMArena Expert1442LMArena2026-10-08
LMArena Expert1463#58 of 273, top 22%LMArena2026-10-08
LMArena Expert1369no-thinkingLMArena2026-10-08

Multilingual

Qwen3 235B-A22B Multilingual benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Non-English1409#88 of 297, top 30%LMArena2026-10-08
LMArena Non-English1397LMArena2026-10-08
LMArena Non-English1384no-thinkingLMArena2026-10-08
LMArena Chinese1465LMArena2026-10-08
LMArena Chinese1481#62 of 285, top 22%LMArena2026-10-08
LMArena Chinese1425no-thinkingLMArena2026-10-08
LMArena French1445#80 of 223, top 36%LMArena2026-10-08
LMArena French1369no-thinkingLMArena2026-10-08
LMArena German1387LMArena2026-10-08
LMArena German1408LMArena2026-10-08
LMArena German1433#62 of 231, top 27%LMArena2026-10-08
LMArena Japanese1386LMArena2026-10-08
LMArena Japanese1399#60 of 211, top 29%LMArena2026-10-08
LMArena Japanese1371no-thinkingLMArena2026-10-08
LMArena Korean1391#68 of 213, top 32%LMArena2026-10-08
LMArena Korean1375LMArena2026-10-08
LMArena Korean1358no-thinkingLMArena2026-10-08
LMArena Russian1400LMArena2026-10-08
LMArena Russian1411#94 of 283, top 34%LMArena2026-10-08
LMArena Russian1394no-thinkingLMArena2026-10-08
LMArena Spanish1389LMArena2026-10-08
LMArena Spanish1430#83 of 226, top 37%LMArena2026-10-08
LMArena Spanish1403no-thinkingLMArena2026-10-08

Instruction Following

Qwen3 235B-A22B Instruction Following benchmark results
BenchmarkScorePositionSettingSourceDate
IFEval81.6%HELM Capabilities
IFEval83.5%#28 of 57, top 50%HELM Capabilities
LMArena Instruction Following1386LMArena2026-10-08
LMArena Instruction Following1408#90 of 298, top 31%LMArena2026-10-08
LMArena Instruction Following1362no-thinkingLMArena2026-10-08

Long Context

Qwen3 235B-A22B Long Context benchmark results
BenchmarkScorePositionSettingSourceDate
Fiction.LiveBench75%#15 of 47, top 32%Epoch AI
Fiction.LiveBench67.7%Epoch AI
Fiction.LiveBench52.9%Epoch AI
LMArena Longer Query1400LMArena2026-10-08
LMArena Longer Query1426#84 of 291, top 29%LMArena2026-10-08
LMArena Longer Query1389no-thinkingLMArena2026-10-08

Writing & Preference

Qwen3 235B-A22B Writing & Preference benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Text1415LMArena2026-10-08
LMArena Text1419#96 of 297, top 33%LMArena2026-10-08
LMArena Text1394no-thinkingLMArena2026-10-08
LMArena Creative Writing1384#102 of 295, top 35%LMArena2026-10-08
LMArena Creative Writing1375LMArena2026-10-08
LMArena Creative Writing1356no-thinkingLMArena2026-10-08
Short-Story Creative Writing82.4%Epoch AI
Short-Story Creative Writing83%#10 of 39, top 26%Epoch AI
EQ-Bench Creative Writing1366#73 of 115, top 64%EQ-Bench
WildBench86.6%Best of 57HELM Capabilities
WildBench82.8%HELM Capabilities
LMArena Multi-Turn1432#81 of 295, top 28%LMArena2026-10-08
LMArena Multi-Turn1406LMArena2026-10-08
LMArena Multi-Turn1398no-thinkingLMArena2026-10-08

API pricing by provider

Qwen3 235B-A22B API prices
RouteInput $/MOutput $/MCached input $/MChecked
alibaba$0.70$2.80—2026-10-10
bedrock$0.22$0.88—2026-10-10
deepinfra$0.09$0.55—2026-10-10
openrouter$0.09$0.55—2026-10-10
together$0.20$0.60—2026-10-10
vertex$0.22$0.88—2026-10-10

Compare Qwen3 235B-A22B

Other Alibaba (Qwen) models

Frequently asked questions

How good is Qwen3 235B-A22B?

Qwen3 235B-A22B by Alibaba (Qwen) ranks 91st of 354 ranked models on the Noometry Index as of October 2026, with a score of 43.5. Its strongest category is long context, where it ranks 26th. API pricing starts at $0.70 per million input tokens and $2.80 per million output tokens, with a 131K-token context window.

How much does Qwen3 235B-A22B cost?

Qwen3 235B-A22B costs $0.70 per million input tokens and $2.80 per million output tokens on Alibaba (Qwen)'s own API.

What is Qwen3 235B-A22B's context window?

Qwen3 235B-A22B accepts up to 131K tokens of input and can write up to 16K tokens in one response.

Is Qwen3 235B-A22B open source?

Yes. Qwen3 235B-A22B's weights are downloadable from Hugging Face (Qwen/Qwen3-235B-A22B-Instruct-2507); check the license for commercial terms.

How fast is Qwen3 235B-A22B?

Qwen3 235B-A22B generated about 85 output tokens per second in the Kagi LLM Benchmark's timed runs. Speed varies by provider, load and reasoning effort.

What are Qwen3 235B-A22B's strengths and weaknesses?

Relative to other ranked models, Qwen3 235B-A22B places best in long context, math, coding and lowest in reasoning, instruction following, writing & preference.

What is Qwen3 235B-A22B best at?

Its best category is long context, where it ranks 26th on Noometry.