Alibaba (Qwen), open weights
Qwen3 14B
Qwen3 14B by Alibaba (Qwen) ranks 225th of 354 ranked models on the Noometry Index as of October 2026, with a score of 35.5. Its strongest category is agentic & tool use, where it ranks 83rd. API pricing starts at $0.35 per million input tokens and $1.40 per million output tokens, with a 131K-token context window.
Last verified
Specifications
- Noometry rank
- #225 of 354
- Index score
- 35.5
- Evidence
- Confirmed 12 results
- Provider
Alibaba (Qwen)
- Released
- April 1, 2025
- Weights
- Open weights
- Reasoning
- Yes
- Context window
- 131K
- Max output
- 8K
- Input price
- $0.35 / M
- Output price
- $1.40 / M
- Blended price
- $0.61 / M
- Output speed
- 79 tokens/s Kagi
- Value
- #91 of 219
- Knowledge cutoff
- April 2025
- Input
- text
- Hugging Face
- Qwen/Qwen3-14B
Category scores
Each category score combines every public result we have in that category.
- Coding 37.3
- Agentic & Tool Use 29.6
- Reasoning 18.5
- Math 38.6
- Knowledge 39.3
- Long Context 38.1
| Category | Score | Rank | Results |
|---|---|---|---|
| Coding | 37.3 | #195 | 1 |
| Agentic & Tool Use | 29.6 | #83 | 1 |
| Reasoning | 18.5 | #280 | 5 |
| Math | 38.6 | #133 | 1 |
| Knowledge | 39.3 | #134 | 2 |
| Long Context | 38.1 | #204 | 1 |
Strengths and weaknesses
Categories where Qwen3 14B places highest and lowest among the models ranked in each, with its score against that category's median.
Strongest categories
| Category | Score | vs median | Rank |
|---|---|---|---|
| Math | 38.6 | +2.1 | #133 of 327, top 41% |
| Knowledge | 39.3 | +1.9 | #134 of 314, top 43% |
| Agentic & Tool Use | 29.6 | −0.7 | #83 of 154, top 54% |
Weakest categories
| Category | Score | vs median | Rank |
|---|---|---|---|
| Reasoning | 18.5 | −5.1 | #280 of 350, top 80% |
| Long Context | 38.1 | −2.8 | #204 of 296, top 69% |
| Coding | 37.3 | −1.5 | #195 of 340, top 58% |
Closest competitors
The models ranked just above and below Qwen3 14B. When scores are this close, price and speed are often the better way to choose.
| Model | Rank | Score | Blended $/M | Speed | |
|---|---|---|---|---|---|
| C4ai Aya Expanse 32b | #221 | 35.9 | — | — | Compare |
| Llama 3.1 Nemotron 51b Instruct | #222 | 35.9 | — | — | Compare |
| Nemotron 4 340b Instruct | #223 | 35.9 | — | — | Compare |
| Llama 3.1 Tulu 3 8b | #224 | 35.7 | — | — | Compare |
| DeepSeek-R1-Distill-Qwen-32B | #226 | 35.5 | — | — | Compare |
| Magistral Medium | #227 | 35.2 | $2.75 | 0 | Compare |
| Gemini 2.0 Flash (Feb 2025) | #228 | 35.1 | — | 92 | Compare |
| C4ai Aya Expanse 8b | #229 | 34.9 | — | — | Compare |
Sponsored placements are available on pages like this one. Advertise on Noometry
Benchmark results
Every published result we track, with its source. Bold rows are the ones used for ranking; where several exist we prefer independent runs over self-reported numbers.
Coding
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| SciCode | 31.6% | #105 of 121, top 87% | Epoch AI |
Agentic & Tool Use
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| Berkeley Function Calling Leaderboard | 41% | #23 of 49, top 47% | fc | Berkeley Function Calling Leaderboard |
Reasoning
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| Kagi LLM Benchmark | 49.1% | #67 of 99, top 68% | Kagi LLM Benchmark | ||
| CritPt | 0% | #130 of 134, top 98% | Epoch AI | ||
| Chess Puzzles | 4% | #99 of 129, top 77% | Epoch AI | 2026-08-30 | |
| Chess Puzzles | 0% | none | Epoch AI | 2026-08-30 | |
| DTBench | 64% | #104 of 151, top 69% | Epoch AI | ||
| LMCA | 18.2% | #95 of 125, top 76% | Epoch AI | ||
| Epoch Capabilities Index | 138.23 | #115 of 213, top 54% | Epoch AI | 2025-04-29 |
Math
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| OTIS Mock AIME 2024-2025 | 66.4% | #99 of 173, top 58% | Epoch AI | 2026-08-28 | |
| OTIS Mock AIME 2024-2025 | 25.8% | none | Epoch AI | 2026-08-30 |
Knowledge
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| GPQA Diamond | 63.8% | #111 of 186, top 60% | Epoch AI | 2026-08-28 | |
| GPQA Diamond | 53.4% | none | Epoch AI | 2026-08-28 | |
| Vectara Hallucination Rate (lower is better) | 5.4% | #14 of 96, top 15% | Vectara Hallucination Leaderboard |
Long Context
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| Fiction.LiveBench | 62.5% | #26 of 47, top 56% | Epoch AI |
API pricing by provider
| Route | Input $/M | Output $/M | Cached input $/M | Checked |
|---|---|---|---|---|
| alibaba | $0.35 | $1.40 | — | 2026-10-10 |
| openrouter | $0.12 | $0.24 | — | 2026-10-10 |
Compare Qwen3 14B
- Qwen3 14B vs Qwen2-72B
- Qwen3 14B vs Llama 3.1 Tulu 3 8b
- Qwen3 14B vs DeepSeek-R1-Distill-Qwen-32B
- Qwen3 14B vs Nemotron 4 340b Instruct
- Qwen3 14B vs Magistral Medium
- Qwen3 14B vs Llama 3.1 Nemotron 51b Instruct
- Qwen3 14B vs Gemini 2.0 Flash (Feb 2025)
- Qwen3 14B vs GPT-6 Astra
- Qwen3 14B vs Claude Fable 5.1
- Qwen3 14B vs Gemini 3.8 Flash
- Qwen3 14B vs Kimi K3
- Qwen3 14B vs Grok 4.6
- Qwen3 14B vs GLM-5.3
- Qwen3 14B vs Muse Spark 1.3
Other Alibaba (Qwen) models
- Qwen3.8 Max56.8
- Qwen3.7 Max51.5
- Qwen3.6 Max Preview51.5
- Qwen3.6 Plus47.5
- Qwen3.5 397B-A17B46.0
- Qwen3.8 27B46.0
- Qwen3.5 Max Preview45.3
- Qwen3.7 Plus45.3
Frequently asked questions
How good is Qwen3 14B?
Qwen3 14B by Alibaba (Qwen) ranks 225th of 354 ranked models on the Noometry Index as of October 2026, with a score of 35.5. Its strongest category is agentic & tool use, where it ranks 83rd. API pricing starts at $0.35 per million input tokens and $1.40 per million output tokens, with a 131K-token context window.
How much does Qwen3 14B cost?
Qwen3 14B costs $0.35 per million input tokens and $1.40 per million output tokens on Alibaba (Qwen)'s own API.
What is Qwen3 14B's context window?
Qwen3 14B accepts up to 131K tokens of input and can write up to 8K tokens in one response.
Is Qwen3 14B open source?
Yes. Qwen3 14B's weights are downloadable from Hugging Face (Qwen/Qwen3-14B); check the license for commercial terms.
How fast is Qwen3 14B?
Qwen3 14B generated about 79 output tokens per second in the Kagi LLM Benchmark's timed runs. Speed varies by provider, load and reasoning effort.
What are Qwen3 14B's strengths and weaknesses?
Relative to other ranked models, Qwen3 14B places best in math, knowledge, agentic & tool use and lowest in reasoning, long context, coding.
What is Qwen3 14B best at?
Its best category is agentic & tool use, where it ranks 83rd on Noometry.