Alibaba (Qwen), open weights

# Qwen3-4B

> Qwen3-4B by Alibaba (Qwen), released April 2025. Ranked #264 of 354 with a Noometry Index of 31.9. Scores, sources and comparisons.
- Canonical page: https://noometry.com/models/qwen3-4b
- Last updated: 2026-10-10
- Title: Qwen3-4B Benchmarks, Price & Rank (October 2026) | Noometry

Qwen3-4B by Alibaba (Qwen) ranks 264th of 354 ranked models on the Noometry Index as of October 2026, with a score of 31.9. Its strongest category is agentic & tool use, where it ranks 100th.

Last verified October 10, 2026

## Specifications

- **Noometry rank:** #264 of 354
- **Index score:** 31.9
- **Evidence:** Confirmed 6 results
- **Provider:** [![](/logos/alibaba.svg) Alibaba (Qwen)](https://noometry.com/providers/alibaba)
- **Released:** April 29, 2025
- **Weights:** Open weights
- **Reasoning:** Unknown
- **Context window:** —
- **Max output:** —
- **Input price:** Not listed
- **Output price:** Not listed
- **Blended price:** Not listed
- **Output speed:** Not measured
- **Value:** Not ranked
- **Knowledge cutoff:** Unknown

## Category scores

Each category score combines every public result we have in that category.

Qwen3-4B category scores

1.  Agentic & Tool Use 27.6
2.  Reasoning 19.2
3.  Math 29.7
4.  Knowledge 33.0
5.  1520253035

Qwen3-4B category ranks
| Category | Score | Rank | Results |
| --- | --- | --- | --- |
| [Agentic & Tool Use](https://noometry.com/best/agentic) | 27.6 | #100 | 1 |
| [Reasoning](https://noometry.com/best/reasoning) | 19.2 | #268 | 1 |
| [Math](https://noometry.com/best/math) | 29.7 | #240 | 2 |
| [Knowledge](https://noometry.com/best/knowledge) | 33.0 | #208 | 2 |

## Strengths and weaknesses

Categories where Qwen3-4B places highest and lowest among the models ranked in each, with its score against that category's median.

### Strongest categories

Qwen3-4B: strongest categories
| Category | Score | vs median | Rank |
| --- | --- | --- | --- |
| [Agentic & Tool Use](https://noometry.com/best/agentic) | 27.6 | −2.8 | #100 of 154, top 65% |
| [Knowledge](https://noometry.com/best/knowledge) | 33.0 | −4.3 | #208 of 314, top 67% |

### Weakest categories

Qwen3-4B: weakest categories
| Category | Score | vs median | Rank |
| --- | --- | --- | --- |
| [Reasoning](https://noometry.com/best/reasoning) | 19.2 | −4.4 | #268 of 350, top 77% |
| [Math](https://noometry.com/best/math) | 29.7 | −6.9 | #240 of 327, top 74% |

## Closest competitors

The models ranked just above and below Qwen3-4B. When scores are this close, price and speed are often the better way to choose.

Models ranked closest to Qwen3-4B
| Model | Rank | Score | Blended $/M | Speed |  |
| --- | --- | --- | --- | --- | --- |
| [Falcon-180B](https://noometry.com/models/falcon-180b) | #260 | 32.2 | — | — | [Compare](https://noometry.com/compare/falcon-180b-vs-qwen3-4b) |
| [Gemini 1.5 Pro (May 2024)](https://noometry.com/models/gemini-1-5-pro) | #261 | 32.1 | — | — | [Compare](https://noometry.com/compare/gemini-1-5-pro-vs-qwen3-4b) |
| [Gemma 3 12B](https://noometry.com/models/gemma-3-12b) | #262 | 32.1 | $0.075 | — | [Compare](https://noometry.com/compare/gemma-3-12b-vs-qwen3-4b) |
| [Mistral Large](https://noometry.com/models/mistral-large) | #263 | 31.9 | $3 | — | [Compare](https://noometry.com/compare/mistral-large-vs-qwen3-4b) |
| [Amazon Nova Lite](https://noometry.com/models/amazon-nova-lite) | #265 | 31.9 | $0.10 | — | [Compare](https://noometry.com/compare/amazon-nova-lite-vs-qwen3-4b) |
| [Mistral Medium 3.1](https://noometry.com/models/mistral-medium-3-1) | #266 | 31.9 | $0.80 | — | [Compare](https://noometry.com/compare/mistral-medium-3-1-vs-qwen3-4b) |
| [Qwen2.5 72B Instruct](https://noometry.com/models/qwen2-5-72b-instruct) | #267 | 31.9 | $2.45 | — | [Compare](https://noometry.com/compare/qwen2-5-72b-instruct-vs-qwen3-4b) |
| [Llama2 70b Steerlm Chat](https://noometry.com/models/llama2-70b-steerlm-chat) | #268 | 31.8 | — | — | [Compare](https://noometry.com/compare/llama2-70b-steerlm-chat-vs-qwen3-4b) |

Sponsored placements are available on pages like this one. [Advertise on Noometry](https://noometry.com/advertise)

## Benchmark results

Every published result we track, with its source. Bold rows are the ones used for ranking; where several exist we prefer independent runs over self-reported numbers.

### Agentic & Tool Use

Qwen3-4B Agentic & Tool Use benchmark results
| Benchmark | Score | Position | Setting | Source | Date |
| --- | --- | --- | --- | --- | --- |
| [Berkeley Function Calling Leaderboard](https://noometry.com/benchmarks/bfcl) | 35.7% | #29 of 49, top 60% | fc | [Berkeley Function Calling Leaderboard](https://gorilla.cs.berkeley.edu/leaderboard.html) |  |

### Reasoning

Qwen3-4B Reasoning benchmark results
| Benchmark | Score | Position | Setting | Source | Date |
| --- | --- | --- | --- | --- | --- |
| [Chess Puzzles](https://noometry.com/benchmarks/chess-puzzles) | 4% | #100 of 129, top 78% |  | [Epoch AI](https://epoch.ai/benchmarks) | 2026-08-28 |
| [Chess Puzzles](https://noometry.com/benchmarks/chess-puzzles) | 1% |  | none | [Epoch AI](https://epoch.ai/benchmarks) | 2026-08-28 |

### Math

Qwen3-4B Math benchmark results
| Benchmark | Score | Position | Setting | Source | Date |
| --- | --- | --- | --- | --- | --- |
| [MathArena Final-Answer Competitions](https://noometry.com/benchmarks/matharena) | 38.5% | #29 of 29, top 100% |  | [MathArena](https://matharena.ai/) |  |
| [OTIS Mock AIME 2024-2025](https://noometry.com/benchmarks/otis-mock-aime) | 52.2% | #111 of 173, top 65% |  | [Epoch AI](https://epoch.ai/benchmarks) | 2026-08-28 |

### Knowledge

Qwen3-4B Knowledge benchmark results
| Benchmark | Score | Position | Setting | Source | Date |
| --- | --- | --- | --- | --- | --- |
| [GPQA Diamond](https://noometry.com/benchmarks/gpqa-diamond) | 52.3% | #124 of 186, top 67% |  | [Epoch AI](https://epoch.ai/benchmarks) | 2026-08-28 |
| [GPQA Diamond](https://noometry.com/benchmarks/gpqa-diamond) | 45.8% |  |  | [Epoch AI](https://epoch.ai/benchmarks) | 2026-08-28 |
| [GPQA Diamond](https://noometry.com/benchmarks/gpqa-diamond) | 43.6% |  | none | [Epoch AI](https://epoch.ai/benchmarks) | 2026-08-28 |
| [Vectara Hallucination Rate](https://noometry.com/benchmarks/vectara-hallucination) (lower is better) | 5.7% | #19 of 96, top 20% |  | [Vectara Hallucination Leaderboard](https://github.com/vectara/hallucination-leaderboard) |  |

## Compare Qwen3-4B

-   [Qwen3-4B vs Qwen3-30B-A3B](https://noometry.com/compare/qwen3-30b-a3b-vs-qwen3-4b)
-   [Qwen3-4B vs Mistral Large](https://noometry.com/compare/mistral-large-vs-qwen3-4b)
-   [Qwen3-4B vs Amazon Nova Lite](https://noometry.com/compare/amazon-nova-lite-vs-qwen3-4b)
-   [Qwen3-4B vs Gemma 3 12B](https://noometry.com/compare/gemma-3-12b-vs-qwen3-4b)
-   [Qwen3-4B vs Mistral Medium 3.1](https://noometry.com/compare/mistral-medium-3-1-vs-qwen3-4b)
-   [Qwen3-4B vs Gemini 1.5 Pro (May 2024)](https://noometry.com/compare/gemini-1-5-pro-vs-qwen3-4b)
-   [Qwen3-4B vs Qwen2.5 72B Instruct](https://noometry.com/compare/qwen2-5-72b-instruct-vs-qwen3-4b)
-   [Qwen3-4B vs GPT-6 Astra](https://noometry.com/compare/gpt-6-astra-vs-qwen3-4b)
-   [Qwen3-4B vs Claude Fable 5.1](https://noometry.com/compare/claude-fable-5-1-vs-qwen3-4b)
-   [Qwen3-4B vs Gemini 3.8 Flash](https://noometry.com/compare/gemini-3-8-flash-vs-qwen3-4b)
-   [Qwen3-4B vs Kimi K3](https://noometry.com/compare/kimi-k3-vs-qwen3-4b)
-   [Qwen3-4B vs Grok 4.6](https://noometry.com/compare/grok-4-6-vs-qwen3-4b)
-   [Qwen3-4B vs GLM-5.3](https://noometry.com/compare/glm-5-3-vs-qwen3-4b)
-   [Qwen3-4B vs Muse Spark 1.3](https://noometry.com/compare/muse-spark-1-3-vs-qwen3-4b)

## Other Alibaba (Qwen) models

-   [Qwen3.8 Max](https://noometry.com/models/qwen3-8-max)56.8
-   [Qwen3.7 Max](https://noometry.com/models/qwen3-7-max)51.5
-   [Qwen3.6 Max Preview](https://noometry.com/models/qwen3-6-max-preview)51.5
-   [Qwen3.6 Plus](https://noometry.com/models/qwen3-6-plus)47.5
-   [Qwen3.5 397B-A17B](https://noometry.com/models/qwen3-5-397b-a17b)46.0
-   [Qwen3.8 27B](https://noometry.com/models/qwen3-8-27b)46.0
-   [Qwen3.5 Max Preview](https://noometry.com/models/qwen3-5-max-preview)45.3
-   [Qwen3.7 Plus](https://noometry.com/models/qwen3-7-plus)45.3

## Frequently asked questions

### How good is Qwen3-4B?

Qwen3-4B by Alibaba (Qwen) ranks 264th of 354 ranked models on the Noometry Index as of October 2026, with a score of 31.9. Its strongest category is agentic & tool use, where it ranks 100th.

### Is Qwen3-4B open source?

Yes. Qwen3-4B's weights are downloadable; check the license for commercial terms.

### What are Qwen3-4B's strengths and weaknesses?

Relative to other ranked models, Qwen3-4B places best in agentic & tool use, knowledge and lowest in reasoning, math.

### What is Qwen3-4B best at?

Its best category is agentic & tool use, where it ranks 100th on Noometry.

### Cite this page

Noometry. (2026). Qwen3-4B benchmarks and pricing. Retrieved October 10, 2026, from https://noometry.com/models/qwen3-4b

Quote Noometry with a link back to this page. It is also available in [Markdown](https://noometry.com/md/models/qwen3-4b.md).
