Google, open weights
Gemma 3 4B
Gemma 3 4B by Google ranks 326th of 354 ranked models on the Noometry Index as of October 2026, with a score of 28.1. Its strongest category is agentic & tool use, where it ranks 142nd. API pricing starts at $0.04 per million input tokens and $0.08 per million output tokens, with a 131K-token context window.
Last verified
Specifications
- Noometry rank
- #326 of 354
- Index score
- 28.1
- Evidence
- Confirmed 22 results
- Provider
Google
- Released
- March 12, 2025
- Weights
- Open weights
- Reasoning
- No
- Context window
- 131K
- Max output
- 4K
- Input price
- $0.04 / M
- Output price
- $0.08 / M
- Blended price
- $0.05 / M
- Output speed
- 72 tokens/s Kagi
- Value
- #4 of 219
- Knowledge cutoff
- August 2024
- Input
- text, image
- Hugging Face
- google/gemma-3-4b-it
Category scores
Each category score combines every public result we have in that category.
- Coding 35.9
- Agentic & Tool Use 20.9
- Reasoning 13.2
- Math 16.8
- Knowledge 11.8
- Multilingual 42.5
- Instruction Following 65.2
- Long Context 38.7
- Writing & Preference 42.0
| Category | Score | Rank | Results |
|---|---|---|---|
| Coding | 35.9 | #215 | 1 |
| Agentic & Tool Use | 20.9 | #142 | 1 |
| Reasoning | 13.2 | #335 | 5 |
| Math | 16.8 | #292 | 2 |
| Knowledge | 11.8 | #299 | 3 |
| Multilingual | 42.5 | #194 | 1 |
| Instruction Following | 65.2 | #225 | 1 |
| Long Context | 38.7 | #194 | 1 |
| Writing & Preference | 42.0 | #239 | 4 |
Strengths and weaknesses
Categories where Gemma 3 4B places highest and lowest among the models ranked in each, with its score against that category's median.
Strongest categories
| Category | Score | vs median | Rank |
|---|---|---|---|
| Coding | 35.9 | −2.9 | #215 of 340, top 64% |
| Multilingual | 42.5 | −4.9 | #194 of 297, top 66% |
| Long Context | 38.7 | −2.3 | #194 of 296, top 66% |
Weakest categories
| Category | Score | vs median | Rank |
|---|---|---|---|
| Reasoning | 13.2 | −10.4 | #335 of 350, top 96% |
| Knowledge | 11.8 | −25.5 | #299 of 314, top 96% |
| Agentic & Tool Use | 20.9 | −9.5 | #142 of 154, top 93% |
Closest competitors
The models ranked just above and below Gemma 3 4B. When scores are this close, price and speed are often the better way to choose.
| Model | Rank | Score | Blended $/M | Speed | |
|---|---|---|---|---|---|
| Qwen1.5 4b Chat | #322 | 28.8 | — | — | Compare |
| Llama 3-70B | #323 | 28.8 | — | 104 | Compare |
| GPT-4o | #324 | 28.6 | $4.38 | — | Compare |
| Ministral 8B | #325 | 28.2 | $0.15 | — | Compare |
| GPT-4.1 nano | #327 | 27.9 | $0.18 | 135 | Compare |
| Phi 3 Mini 4k Instruct | #328 | 27.9 | — | — | Compare |
| Yi-34B | #329 | 27.8 | — | — | Compare |
| Llama 4 Scout | #330 | 27.7 | $0.15 | 272 | Compare |
Sponsored placements are available on pages like this one. Advertise on Noometry
Benchmark results
Every published result we track, with its source. Bold rows are the ones used for ranking; where several exist we prefer independent runs over self-reported numbers.
Coding
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Coding | 1230 | #231 of 294, top 79% | LMArena | 2026-10-08 |
Agentic & Tool Use
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| Berkeley Function Calling Leaderboard | 19.6% | #45 of 49, top 92% | prompt | Berkeley Function Calling Leaderboard |
Reasoning
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| Kagi LLM Benchmark | 25.2% | #96 of 99, top 97% | Kagi LLM Benchmark | ||
| Chess Puzzles | 0% | #113 of 129, top 88% | Epoch AI | 2026-08-28 | |
| LMArena Hard Prompts | 1253 | #218 of 297, top 74% | LMArena | 2026-10-08 | |
| DTBench | 50.9% | #134 of 151, top 89% | Epoch AI | ||
| LMCA | 2.8% | #125 of 125, top 100% | Epoch AI | ||
| Epoch Capabilities Index | 116.02 | #184 of 213, top 87% | Epoch AI | 2025-03-12 |
Math
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| OTIS Mock AIME 2024-2025 | 7.5% | #139 of 173, top 81% | Epoch AI | 2026-08-28 | |
| LMArena Math | 1239 | #221 of 285, top 78% | LMArena | 2026-10-08 |
Knowledge
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| GPQA Diamond | 23.2% | #183 of 186, top 99% | Epoch AI | 2026-08-28 | |
| Vectara Hallucination Rate (lower is better) | 6.4% | #24 of 96, top 25% | Vectara Hallucination Leaderboard | ||
| LMArena Expert | 1223 | #212 of 273, top 78% | LMArena | 2026-10-08 |
Multilingual
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Non-English | 1273 | #194 of 297, top 66% | LMArena | 2026-10-08 | |
| LMArena German | 1281 | #160 of 231, top 70% | LMArena | 2026-10-08 | |
| LMArena Russian | 1294 | #180 of 283, top 64% | LMArena | 2026-10-08 |
Instruction Following
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Instruction Following | 1239 | #223 of 298, top 75% | LMArena | 2026-10-08 |
Long Context
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Longer Query | 1273 | #208 of 291, top 72% | LMArena | 2026-10-08 |
Writing & Preference
| Benchmark | Score | Position | Setting | Source | Date |
|---|---|---|---|---|---|
| LMArena Text | 1291 | #196 of 297, top 66% | LMArena | 2026-10-08 | |
| LMArena Creative Writing | 1271 | #188 of 295, top 64% | LMArena | 2026-10-08 | |
| EQ-Bench Creative Writing | 1068 | #94 of 115, top 82% | EQ-Bench | ||
| LMArena Multi-Turn | 1255 | #220 of 295, top 75% | LMArena | 2026-10-08 |
API pricing by provider
| Route | Input $/M | Output $/M | Cached input $/M | Checked |
|---|---|---|---|---|
| bedrock | $0.04 | $0.08 | — | 2026-10-10 |
| deepinfra | $0.05 | $0.10 | — | 2026-10-10 |
| openrouter | $0.05 | $0.10 | — | 2026-10-10 |
Compare Gemma 3 4B
- Gemma 3 4B vs Gemma 3 27B
- Gemma 3 4B vs Ministral 8B
- Gemma 3 4B vs GPT-4.1 nano
- Gemma 3 4B vs GPT-4o
- Gemma 3 4B vs Phi 3 Mini 4k Instruct
- Gemma 3 4B vs Llama 3-70B
- Gemma 3 4B vs Yi-34B
- Gemma 3 4B vs GPT-6 Astra
- Gemma 3 4B vs Claude Fable 5.1
- Gemma 3 4B vs Kimi K3
- Gemma 3 4B vs Grok 4.6
- Gemma 3 4B vs Qwen3.8 Max
- Gemma 3 4B vs GLM-5.3
- Gemma 3 4B vs Muse Spark 1.3
Other Google models
Frequently asked questions
How good is Gemma 3 4B?
Gemma 3 4B by Google ranks 326th of 354 ranked models on the Noometry Index as of October 2026, with a score of 28.1. Its strongest category is agentic & tool use, where it ranks 142nd. API pricing starts at $0.04 per million input tokens and $0.08 per million output tokens, with a 131K-token context window.
How much does Gemma 3 4B cost?
Gemma 3 4B costs $0.04 per million input tokens and $0.08 per million output tokens on bedrock.
What is Gemma 3 4B's context window?
Gemma 3 4B accepts up to 131K tokens of input and can write up to 4K tokens in one response.
Is Gemma 3 4B open source?
Yes. Gemma 3 4B's weights are downloadable from Hugging Face (google/gemma-3-4b-it); check the license for commercial terms.
How fast is Gemma 3 4B?
Gemma 3 4B generated about 72 output tokens per second in the Kagi LLM Benchmark's timed runs. Speed varies by provider, load and reasoning effort.
What are Gemma 3 4B's strengths and weaknesses?
Relative to other ranked models, Gemma 3 4B places best in coding, multilingual, long context and lowest in reasoning, knowledge, agentic & tool use.
What is Gemma 3 4B best at?
Its best category is agentic & tool use, where it ranks 142nd on Noometry.