Alibaba (Qwen), open weights
Qwen3-4B
Qwen3-4B by Alibaba (Qwen) ranks 264th of 354 ranked models on the Noometry Index as of October 2026, with a score of 31.9. Its strongest category is agentic & tool use, where it ranks 100th.
Last verified
Specifications
- Noometry rank
- #264 of 354
- Index score
- 31.9
- Evidence
- Confirmed 6 results
- Provider
Alibaba (Qwen)
- Released
- April 29, 2025
- Weights
- Open weights
- Reasoning
- Unknown
- Context window
- —
- Max output
- —
- Input price
- Not listed
- Output price
- Not listed
- Blended price
- Not listed
- Output speed
- Not measured
- Value
- Not ranked
- Knowledge cutoff
- Unknown
Category scores
Each category score combines every public result we have in that category.
- Agentic & Tool Use 27.6
- Reasoning 19.2
- Math 29.7
- Knowledge 33.0
| Category | Score | Rank | Results |
|---|---|---|---|
| Agentic & Tool Use | 27.6 | #100 | 1 |
| Reasoning | 19.2 | #268 | 1 |
| Math | 29.7 | #240 | 2 |
| Knowledge | 33.0 | #208 | 2 |
Strengths and weaknesses
Categories where Qwen3-4B places highest and lowest among the models ranked in each, with its score against that category's median.
Strongest categories
| Category | Score | vs median | Rank |
|---|---|---|---|
| Agentic & Tool Use | 27.6 | −2.8 | #100 of 154, top 65% |
| Knowledge | 33.0 | −4.3 | #208 of 314, top 67% |
Closest competitors
The models ranked just above and below Qwen3-4B. When scores are this close, price and speed are often the better way to choose.
| Model | Rank | Score | Blended $/M | Speed | |
|---|---|---|---|---|---|
| Falcon-180B | #260 | 32.2 | — | — | Compare |
| Gemini 1.5 Pro (May 2024) | #261 | 32.1 | — | — | Compare |
| Gemma 3 12B | #262 | 32.1 | $0.075 | — | Compare |
| Mistral Large | #263 | 31.9 | $3 | — | Compare |
| Amazon Nova Lite | #265 | 31.9 | $0.10 | — | Compare |
| Mistral Medium 3.1 | #266 | 31.9 | $0.80 | — | Compare |
| Qwen2.5 72B Instruct | #267 | 31.9 | $2.45 | — | Compare |
| Llama2 70b Steerlm Chat | #268 | 31.8 | — | — | Compare |
Sponsored placements are available on pages like this one. Advertise on Noometry
Benchmark results
Every published result we track, with its source. Bold rows are the ones used for ranking; where several exist we prefer independent runs over self-reported numbers.
Agentic & Tool Use
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| Berkeley Function Calling Leaderboard | 35.7% | #29 of 49, top 60% | fc | Berkeley Function Calling Leaderboard |
Reasoning
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| Chess Puzzles | 4% | #100 of 129, top 78% | Epoch AI | 2026-08-28 | |
| Chess Puzzles | 1% | none | Epoch AI | 2026-08-28 |
Math
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| MathArena Final-Answer Competitions | 38.5% | #29 of 29, top 100% | MathArena | ||
| OTIS Mock AIME 2024-2025 | 52.2% | #111 of 173, top 65% | Epoch AI | 2026-08-28 |
Knowledge
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| GPQA Diamond | 52.3% | #124 of 186, top 67% | Epoch AI | 2026-08-28 | |
| GPQA Diamond | 45.8% | Epoch AI | 2026-08-28 | ||
| GPQA Diamond | 43.6% | none | Epoch AI | 2026-08-28 | |
| Vectara Hallucination Rate (lower is better) | 5.7% | #19 of 96, top 20% | Vectara Hallucination Leaderboard |
Compare Qwen3-4B
- Qwen3-4B vs Qwen3-30B-A3B
- Qwen3-4B vs Mistral Large
- Qwen3-4B vs Amazon Nova Lite
- Qwen3-4B vs Gemma 3 12B
- Qwen3-4B vs Mistral Medium 3.1
- Qwen3-4B vs Gemini 1.5 Pro (May 2024)
- Qwen3-4B vs Qwen2.5 72B Instruct
- Qwen3-4B vs GPT-6 Astra
- Qwen3-4B vs Claude Fable 5.1
- Qwen3-4B vs Gemini 3.8 Flash
- Qwen3-4B vs Kimi K3
- Qwen3-4B vs Grok 4.6
- Qwen3-4B vs GLM-5.3
- Qwen3-4B vs Muse Spark 1.3
Other Alibaba (Qwen) models
- Qwen3.8 Max56.8
- Qwen3.7 Max51.5
- Qwen3.6 Max Preview51.5
- Qwen3.6 Plus47.5
- Qwen3.5 397B-A17B46.0
- Qwen3.8 27B46.0
- Qwen3.5 Max Preview45.3
- Qwen3.7 Plus45.3
Frequently asked questions
How good is Qwen3-4B?
Qwen3-4B by Alibaba (Qwen) ranks 264th of 354 ranked models on the Noometry Index as of October 2026, with a score of 31.9. Its strongest category is agentic & tool use, where it ranks 100th.
Is Qwen3-4B open source?
Yes. Qwen3-4B's weights are downloadable; check the license for commercial terms.
What are Qwen3-4B's strengths and weaknesses?
Relative to other ranked models, Qwen3-4B places best in agentic & tool use, knowledge and lowest in reasoning, math.
What is Qwen3-4B best at?
Its best category is agentic & tool use, where it ranks 100th on Noometry.