Google, open weights

Gemma 3 4B

Gemma 3 4B by Google ranks 326th of 354 ranked models on the Noometry Index as of October 2026, with a score of 28.1. Its strongest category is agentic & tool use, where it ranks 142nd. API pricing starts at $0.04 per million input tokens and $0.08 per million output tokens, with a 131K-token context window.

Last verified

Specifications

Noometry rank
#326 of 354
Index score
28.1
Evidence
Confirmed 22 results
Provider
Google
Released
March 12, 2025
Weights
Open weights
Reasoning
No
Context window
131K
Max output
4K
Input price
$0.04 / M
Output price
$0.08 / M
Blended price
$0.05 / M
Output speed
72 tokens/s Kagi
Value
#4 of 219
Knowledge cutoff
August 2024
Input
text, image

Category scores

Each category score combines every public result we have in that category.

Gemma 3 4B category scores
  1. Coding 35.9
  2. Agentic & Tool Use 20.9
  3. Reasoning 13.2
  4. Math 16.8
  5. Knowledge 11.8
  6. Multilingual 42.5
  7. Instruction Following 65.2
  8. Long Context 38.7
  9. Writing & Preference 42.0
Gemma 3 4B category ranks
CategoryScoreRankResults
Coding35.9#2151
Agentic & Tool Use20.9#1421
Reasoning13.2#3355
Math16.8#2922
Knowledge11.8#2993
Multilingual42.5#1941
Instruction Following65.2#2251
Long Context38.7#1941
Writing & Preference42.0#2394

Strengths and weaknesses

Categories where Gemma 3 4B places highest and lowest among the models ranked in each, with its score against that category's median.

Strongest categories

Gemma 3 4B: strongest categories
CategoryScorevs medianRank
Coding35.9−2.9#215 of 340, top 64%
Multilingual42.5−4.9#194 of 297, top 66%
Long Context38.7−2.3#194 of 296, top 66%

Weakest categories

Gemma 3 4B: weakest categories
CategoryScorevs medianRank
Reasoning13.2−10.4#335 of 350, top 96%
Knowledge11.8−25.5#299 of 314, top 96%
Agentic & Tool Use20.9−9.5#142 of 154, top 93%

Closest competitors

The models ranked just above and below Gemma 3 4B. When scores are this close, price and speed are often the better way to choose.

Models ranked closest to Gemma 3 4B
ModelRankScoreBlended $/MSpeed
Qwen1.5 4b Chat#32228.8——Compare
Llama 3-70B#32328.8—104Compare
GPT-4o#32428.6$4.38—Compare
Ministral 8B#32528.2$0.15—Compare
GPT-4.1 nano#32727.9$0.18135Compare
Phi 3 Mini 4k Instruct#32827.9——Compare
Yi-34B#32927.8——Compare
Llama 4 Scout#33027.7$0.15272Compare

Sponsored placements are available on pages like this one. Advertise on Noometry

Benchmark results

Every published result we track, with its source. Bold rows are the ones used for ranking; where several exist we prefer independent runs over self-reported numbers.

Coding

Gemma 3 4B Coding benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Coding1230#231 of 294, top 79%LMArena2026-10-08

Agentic & Tool Use

Gemma 3 4B Agentic & Tool Use benchmark results
BenchmarkScorePositionSettingSourceDate
Berkeley Function Calling Leaderboard19.6%#45 of 49, top 92%promptBerkeley Function Calling Leaderboard

Reasoning

Gemma 3 4B Reasoning benchmark results
BenchmarkScorePositionSettingSourceDate
Kagi LLM Benchmark25.2%#96 of 99, top 97%Kagi LLM Benchmark
Chess Puzzles0%#113 of 129, top 88%Epoch AI2026-08-28
LMArena Hard Prompts1253#218 of 297, top 74%LMArena2026-10-08
DTBench50.9%#134 of 151, top 89%Epoch AI
LMCA2.8%#125 of 125, top 100%Epoch AI
Epoch Capabilities Index116.02#184 of 213, top 87%Epoch AI2025-03-12

Math

Gemma 3 4B Math benchmark results
BenchmarkScorePositionSettingSourceDate
OTIS Mock AIME 2024-20257.5%#139 of 173, top 81%Epoch AI2026-08-28
LMArena Math1239#221 of 285, top 78%LMArena2026-10-08

Knowledge

Gemma 3 4B Knowledge benchmark results
BenchmarkScorePositionSettingSourceDate
GPQA Diamond23.2%#183 of 186, top 99%Epoch AI2026-08-28
Vectara Hallucination Rate (lower is better)6.4%#24 of 96, top 25%Vectara Hallucination Leaderboard
LMArena Expert1223#212 of 273, top 78%LMArena2026-10-08

Multilingual

Gemma 3 4B Multilingual benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Non-English1273#194 of 297, top 66%LMArena2026-10-08
LMArena German1281#160 of 231, top 70%LMArena2026-10-08
LMArena Russian1294#180 of 283, top 64%LMArena2026-10-08

Instruction Following

Gemma 3 4B Instruction Following benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Instruction Following1239#223 of 298, top 75%LMArena2026-10-08

Long Context

Gemma 3 4B Long Context benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Longer Query1273#208 of 291, top 72%LMArena2026-10-08

Writing & Preference

Gemma 3 4B Writing & Preference benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Text1291#196 of 297, top 66%LMArena2026-10-08
LMArena Creative Writing1271#188 of 295, top 64%LMArena2026-10-08
EQ-Bench Creative Writing1068#94 of 115, top 82%EQ-Bench
LMArena Multi-Turn1255#220 of 295, top 75%LMArena2026-10-08

API pricing by provider

Gemma 3 4B API prices
RouteInput $/MOutput $/MCached input $/MChecked
bedrock$0.04$0.08—2026-10-10
deepinfra$0.05$0.10—2026-10-10
openrouter$0.05$0.10—2026-10-10

Compare Gemma 3 4B

Other Google models

Frequently asked questions

How good is Gemma 3 4B?

Gemma 3 4B by Google ranks 326th of 354 ranked models on the Noometry Index as of October 2026, with a score of 28.1. Its strongest category is agentic & tool use, where it ranks 142nd. API pricing starts at $0.04 per million input tokens and $0.08 per million output tokens, with a 131K-token context window.

How much does Gemma 3 4B cost?

Gemma 3 4B costs $0.04 per million input tokens and $0.08 per million output tokens on bedrock.

What is Gemma 3 4B's context window?

Gemma 3 4B accepts up to 131K tokens of input and can write up to 4K tokens in one response.

Is Gemma 3 4B open source?

Yes. Gemma 3 4B's weights are downloadable from Hugging Face (google/gemma-3-4b-it); check the license for commercial terms.

How fast is Gemma 3 4B?

Gemma 3 4B generated about 72 output tokens per second in the Kagi LLM Benchmark's timed runs. Speed varies by provider, load and reasoning effort.

What are Gemma 3 4B's strengths and weaknesses?

Relative to other ranked models, Gemma 3 4B places best in coding, multilingual, long context and lowest in reasoning, knowledge, agentic & tool use.

What is Gemma 3 4B best at?

Its best category is agentic & tool use, where it ranks 142nd on Noometry.