Meta, open weights

Llama 4 Maverick

Llama 4 Maverick by Meta ranks 282nd of 354 ranked models on the Noometry Index as of October 2026, with a score of 30.9. Its strongest category is agentic & tool use, where it ranks 91st. API pricing starts at $0.19 per million input tokens and $0.65 per million output tokens, with a 128K-token context window.

Last verified

Specifications

Noometry rank
#282 of 354
Index score
30.9
Evidence
Confirmed 54 results
Provider
Meta
Released
April 5, 2025
Weights
Open weights
Reasoning
No
Context window
128K
Max output
4K
Input price
$0.19 / M
Output price
$0.65 / M
Blended price
$0.30 / M
Output speed
456 tokens/s Kagi
Value
#58 of 219
Knowledge cutoff
August 2024
Input
text, image

Category scores

Each category score combines every public result we have in that category.

Llama 4 Maverick category scores
  1. Coding 26.6
  2. Agentic & Tool Use 28.2
  3. Reasoning 10.1
  4. Math 26.0
  5. Knowledge 33.4
  6. Multimodal 31.6
  7. Multilingual 42.2
  8. Instruction Following 71.7
  9. Long Context 31.4
  10. Writing & Preference 38.8
Llama 4 Maverick category ranks
CategoryScoreRankResults
Coding26.6#3247
Agentic & Tool Use28.2#911
Reasoning10.1#34210
Math26.0#2624
Knowledge33.4#2047
Multimodal31.6#1052
Multilingual42.2#1951
Instruction Following71.7#1462
Long Context31.4#2792
Writing & Preference38.8#2526

Strengths and weaknesses

Categories where Llama 4 Maverick places highest and lowest among the models ranked in each, with its score against that category's median.

Strongest categories

Llama 4 Maverick: strongest categories
CategoryScorevs medianRank
Instruction Following71.7+0.4#146 of 305, top 48%
Agentic & Tool Use28.2−2.2#91 of 154, top 60%
Knowledge33.4−3.9#204 of 314, top 65%

Weakest categories

Llama 4 Maverick: weakest categories
CategoryScorevs medianRank
Reasoning10.1−13.5#342 of 350, top 98%
Coding26.6−12.2#324 of 340, top 96%
Long Context31.4−9.6#279 of 296, top 95%

Closest competitors

The models ranked just above and below Llama 4 Maverick. When scores are this close, price and speed are often the better way to choose.

Models ranked closest to Llama 4 Maverick
ModelRankScoreBlended $/MSpeed
Mistral Small 3#27831.2$0.0575—Compare
Phi-4#27931.2$0.0875—Compare
Mistral Small 3.2#28031.2$0.1368Compare
Amazon Nova Pro#28131.0$1.40—Compare
Phi-4 Mini#28330.9$0.13—Compare
Gemma 3 27B#28430.8$0.1062Compare
Qwen1.5-72B#28530.8——Compare
Granite 3.0 2b Instruct#28630.8——Compare

Sponsored placements are available on pages like this one. Advertise on Noometry

Benchmark results

Every published result we track, with its source. Bold rows are the ones used for ranking; where several exist we prefer independent runs over self-reported numbers.

Coding

Llama 4 Maverick Coding benchmark results
BenchmarkScorePositionSettingSourceDate
SWE-bench Verified (bash only)21%#36 of 39, top 93%SWE-bench2025-07-20
Aider Polyglot15.6%#38 of 44, top 87%Epoch AI
SciCode33.1%#103 of 121, top 86%Epoch AI
WeirdML24.5%#101 of 119, top 85%Epoch AI
BigCodeBench Instruct49.7%#3 of 64, top 5%BigCodeBench2025-04-05
LMArena Coding1302#198 of 294, top 68%LMArena2026-10-08
BigCodeBench Complete61.4%#2 of 66, top 4%BigCodeBench2025-04-05
ALE-Bench172.97#104 of 105, top 100%Epoch AI

Agentic & Tool Use

Llama 4 Maverick Agentic & Tool Use benchmark results
BenchmarkScorePositionSettingSourceDate
Berkeley Function Calling Leaderboard37.3%#26 of 49, top 54%fcBerkeley Function Calling Leaderboard

Reasoning

Llama 4 Maverick Reasoning benchmark results
BenchmarkScorePositionSettingSourceDate
ARC-AGI-20%#80 of 83, top 97%Epoch AI
SimpleBench27.7%#61 of 77, top 80%Epoch AI
Kagi LLM Benchmark55.9%#49 of 99, top 50%Kagi LLM Benchmark
NYT Connections (extended)8%#89 of 91, top 98%Lech Mazur benchmarks
ARC-AGI-14.4%#80 of 83, top 97%Epoch AI
CritPt0%#118 of 134, top 89%Epoch AI
CritPt0%#118 of 134, top 89%Epoch AI
EnigmaEval0.6%#37 of 38, top 98%Epoch AI
LMArena Hard Prompts1281#200 of 297, top 68%LMArena2026-10-08
DTBench61.9%#110 of 151, top 73%Epoch AI
LMCA15.9%#103 of 125, top 83%Epoch AI
Epoch Capabilities Index132.2#133 of 213, top 63%Epoch AI2025-04-06
ForecastBench57.5#56 of 72, top 78%Epoch AI

Math

Llama 4 Maverick Math benchmark results
BenchmarkScorePositionSettingSourceDate
OTIS Mock AIME 2024-202520.6%#126 of 173, top 73%Epoch AI2025-04-08
Omni-MATH42.2%#24 of 57, top 43%HELM Capabilities
LMArena Math1299#185 of 285, top 65%LMArena2026-10-08
MATH Level 573%#29 of 79, top 37%Epoch AI2025-04-08
FrontierMath (Feb 2025 set)0.7%#62 of 68, top 92%Epoch AI2025-04-08

Knowledge

Llama 4 Maverick Knowledge benchmark results
BenchmarkScorePositionSettingSourceDate
GPQA Diamond67%#102 of 186, top 55%Epoch AI2025-04-08
Humanity's Last Exam5.7%#33 of 41, top 81%Epoch AI
MMLU-Pro81%#15 of 58, top 26%HELM Capabilities
Confabulations (lower is better)22.6%#38 of 51, top 75%Lech Mazur benchmarks
Vectara Hallucination Rate (lower is better)8.2%#37 of 96, top 39%Vectara Hallucination Leaderboard
GPQA (HELM)65%#19 of 57, top 34%HELM Capabilities
LMArena Expert1259#190 of 273, top 70%LMArena2026-10-08

Multimodal

Llama 4 Maverick Multimodal benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Vision1142#99 of 122, top 82%LMArena2026-10-09
GeoBench52%#20 of 25, top 80%Epoch AI
SpatialViz-Bench31.8%#8 of 8, top 100%Epoch AI

Multilingual

Llama 4 Maverick Multilingual benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Non-English1269#195 of 297, top 66%LMArena2026-10-08
LMArena Chinese1277#194 of 285, top 69%LMArena2026-10-08
LMArena French1259#181 of 223, top 82%LMArena2026-10-08
LMArena German1291#152 of 231, top 66%LMArena2026-10-08
LMArena Japanese1207#155 of 211, top 74%LMArena2026-10-08
LMArena Korean1203#159 of 213, top 75%LMArena2026-10-08
LMArena Russian1286#184 of 283, top 66%LMArena2026-10-08
LMArena Spanish1293#162 of 226, top 72%LMArena2026-10-08

Instruction Following

Llama 4 Maverick Instruction Following benchmark results
BenchmarkScorePositionSettingSourceDate
IFEval90.8%#8 of 57, top 15%HELM Capabilities
LMArena Instruction Following1267#198 of 298, top 67%LMArena2026-10-08

Long Context

Llama 4 Maverick Long Context benchmark results
BenchmarkScorePositionSettingSourceDate
Fiction.LiveBench46.2%#38 of 47, top 81%Epoch AI
LMArena Longer Query1280#203 of 291, top 70%LMArena2026-10-08

Writing & Preference

Llama 4 Maverick Writing & Preference benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Text1287#201 of 297, top 68%LMArena2026-10-08
LMArena Creative Writing1267#194 of 295, top 66%LMArena2026-10-08
Short-Story Creative Writing62%#37 of 39, top 95%Epoch AI
EQ-Bench Creative Writing860#102 of 115, top 89%EQ-Bench
WildBench80%#31 of 57, top 55%HELM Capabilities
LMArena Multi-Turn1289#196 of 295, top 67%LMArena2026-10-08

API pricing by provider

Llama 4 Maverick API prices
RouteInput $/MOutput $/MCached input $/MChecked
azure$0.25$1—2026-10-10
bedrock$0.24$0.97—2026-10-10
deepinfra$0.20$0.80—2026-10-10
openrouter$0.19$0.65$0.052026-10-10
vertex$0.35$1.15—2026-10-10

Compare Llama 4 Maverick

Other Meta models

Frequently asked questions

How good is Llama 4 Maverick?

Llama 4 Maverick by Meta ranks 282nd of 354 ranked models on the Noometry Index as of October 2026, with a score of 30.9. Its strongest category is agentic & tool use, where it ranks 91st. API pricing starts at $0.19 per million input tokens and $0.65 per million output tokens, with a 128K-token context window.

How much does Llama 4 Maverick cost?

Llama 4 Maverick costs $0.19 per million input tokens and $0.65 per million output tokens on openrouter, with cached input at $0.05.

What is Llama 4 Maverick's context window?

Llama 4 Maverick accepts up to 128K tokens of input and can write up to 4K tokens in one response.

Is Llama 4 Maverick open source?

Yes. Llama 4 Maverick's weights are downloadable from Hugging Face (meta-llama/Llama-4-Maverick-17B-128E-Instruct); check the license for commercial terms.

How fast is Llama 4 Maverick?

Llama 4 Maverick generated about 456 output tokens per second in the Kagi LLM Benchmark's timed runs. Speed varies by provider, load and reasoning effort.

What are Llama 4 Maverick's strengths and weaknesses?

Relative to other ranked models, Llama 4 Maverick places best in instruction following, agentic & tool use, knowledge and lowest in reasoning, coding, long context.

What is Llama 4 Maverick best at?

Its best category is agentic & tool use, where it ranks 91st on Noometry.