Alibaba (Qwen), open weights
Qwen3 8B
Qwen3 8B by Alibaba (Qwen) ranks 238th of 354 ranked models on the Noometry Index as of October 2026, with a score of 33.7. Its strongest category is agentic & tool use, where it ranks 78th. API pricing starts at $0.18 per million input tokens and $0.70 per million output tokens, with a 131K-token context window.
Last verified
Specifications
- Noometry rank
- #238 of 354
- Index score
- 33.7
- Evidence
- Confirmed 11 results
- Provider
Alibaba (Qwen)
- Released
- April 1, 2025
- Weights
- Open weights
- Reasoning
- Yes
- Context window
- 131K
- Max output
- 8K
- Input price
- $0.18 / M
- Output price
- $0.70 / M
- Blended price
- $0.31 / M
- Output speed
- Not measured
- Value
- #55 of 219
- Knowledge cutoff
- April 2025
- Input
- text
Category scores
Each category score combines every public result we have in that category.
- Coding 34.0
- Agentic & Tool Use 30.2
- Reasoning 16.6
- Math 34.9
- Knowledge 36.1
- Long Context 37.9
| Category | Score | Rank | Results |
|---|---|---|---|
| Coding | 34.0 | #248 | 1 |
| Agentic & Tool Use | 30.2 | #78 | 1 |
| Reasoning | 16.6 | #303 | 4 |
| Math | 34.9 | #191 | 1 |
| Knowledge | 36.1 | #173 | 2 |
| Long Context | 37.9 | #210 | 1 |
Strengths and weaknesses
Categories where Qwen3 8B places highest and lowest among the models ranked in each, with its score against that category's median.
Strongest categories
| Category | Score | vs median | Rank |
|---|---|---|---|
| Agentic & Tool Use | 30.2 | −0.2 | #78 of 154, top 51% |
| Knowledge | 36.1 | −1.2 | #173 of 314, top 56% |
| Math | 34.9 | −1.7 | #191 of 327, top 59% |
Weakest categories
| Category | Score | vs median | Rank |
|---|---|---|---|
| Reasoning | 16.6 | −7.0 | #303 of 350, top 87% |
| Coding | 34.0 | −4.7 | #248 of 340, top 73% |
| Long Context | 37.9 | −3.0 | #210 of 296, top 71% |
Closest competitors
The models ranked just above and below Qwen3 8B. When scores are this close, price and speed are often the better way to choose.
| Model | Rank | Score | Blended $/M | Speed | |
|---|---|---|---|---|---|
| Qwen1.5-110B | #234 | 34.2 | — | — | Compare |
| o1-mini | #235 | 34.0 | — | — | Compare |
| Qwen3.5-9B | #236 | 33.8 | $0.11 | — | Compare |
| Codellama 70b Instruct | #237 | 33.7 | — | — | Compare |
| Grok-2 (Dec 2024) | #239 | 33.7 | — | — | Compare |
| GPT-4.1 mini | #240 | 33.6 | $0.70 | 86 | Compare |
| GPT-5 Nano | #241 | 33.5 | $0.14 | 4 | Compare |
| Mercury 2.5 | #242 | 33.5 | $0.0675 | — | Compare |
Sponsored placements are available on pages like this one. Advertise on Noometry
Benchmark results
Every published result we track, with its source. Bold rows are the ones used for ranking; where several exist we prefer independent runs over self-reported numbers.
Coding
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| SciCode | 22.6% | #116 of 121, top 96% | Epoch AI |
Agentic & Tool Use
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| Berkeley Function Calling Leaderboard | 42.6% | #21 of 49, top 43% | fc | Berkeley Function Calling Leaderboard |
Reasoning
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| CritPt | 0% | #132 of 134, top 99% | Epoch AI | ||
| Chess Puzzles | 5% | #95 of 129, top 74% | Epoch AI | 2026-08-27 | |
| Chess Puzzles | 0% | none | Epoch AI | 2026-08-27 | |
| DTBench | 59.7% | #117 of 151, top 78% | Epoch AI | ||
| LMCA | 8.8% | #116 of 125, top 93% | Epoch AI | ||
| Epoch Capabilities Index | 136.17 | #121 of 213, top 57% | Epoch AI | 2025-04-28 |
Math
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| OTIS Mock AIME 2024-2025 | 56.1% | #107 of 173, top 62% | Epoch AI | 2026-08-27 | |
| OTIS Mock AIME 2024-2025 | 22.2% | none | Epoch AI | 2026-08-27 |
Knowledge
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| GPQA Diamond | 56.8% | #117 of 186, top 63% | Epoch AI | 2026-08-27 | |
| GPQA Diamond | 47.5% | none | Epoch AI | 2026-08-27 | |
| Vectara Hallucination Rate (lower is better) | 4.8% | #7 of 96, top 8% | Vectara Hallucination Leaderboard |
Long Context
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| Fiction.LiveBench | 62.1% | #27 of 47, top 58% | Epoch AI |
API pricing by provider
| Route | Input $/M | Output $/M | Cached input $/M | Checked |
|---|---|---|---|---|
| alibaba | $0.18 | $0.70 | — | 2026-10-10 |
Compare Qwen3 8B
- Qwen3 8B vs Qwen2-72B
- Qwen3 8B vs Codellama 70b Instruct
- Qwen3 8B vs Grok-2 (Dec 2024)
- Qwen3 8B vs Qwen3.5-9B
- Qwen3 8B vs GPT-4.1 mini
- Qwen3 8B vs o1-mini
- Qwen3 8B vs GPT-5 Nano
- Qwen3 8B vs GPT-6 Astra
- Qwen3 8B vs Claude Fable 5.1
- Qwen3 8B vs Gemini 3.8 Flash
- Qwen3 8B vs Kimi K3
- Qwen3 8B vs Grok 4.6
- Qwen3 8B vs GLM-5.3
- Qwen3 8B vs Muse Spark 1.3
Other Alibaba (Qwen) models
- Qwen3.8 Max56.8
- Qwen3.7 Max51.5
- Qwen3.6 Max Preview51.5
- Qwen3.6 Plus47.5
- Qwen3.5 397B-A17B46.0
- Qwen3.8 27B46.0
- Qwen3.5 Max Preview45.3
- Qwen3.7 Plus45.3
Frequently asked questions
How good is Qwen3 8B?
Qwen3 8B by Alibaba (Qwen) ranks 238th of 354 ranked models on the Noometry Index as of October 2026, with a score of 33.7. Its strongest category is agentic & tool use, where it ranks 78th. API pricing starts at $0.18 per million input tokens and $0.70 per million output tokens, with a 131K-token context window.
How much does Qwen3 8B cost?
Qwen3 8B costs $0.18 per million input tokens and $0.70 per million output tokens on Alibaba (Qwen)'s own API.
What is Qwen3 8B's context window?
Qwen3 8B accepts up to 131K tokens of input and can write up to 8K tokens in one response.
Is Qwen3 8B open source?
Yes. Qwen3 8B's weights are downloadable; check the license for commercial terms.
What are Qwen3 8B's strengths and weaknesses?
Relative to other ranked models, Qwen3 8B places best in agentic & tool use, knowledge, math and lowest in reasoning, coding, long context.
What is Qwen3 8B best at?
Its best category is agentic & tool use, where it ranks 78th on Noometry.