Model comparison

Amazon Nova Micro vs Llama 4 Maverick

Amazon Nova Micro and Llama 4 Maverick score almost the same on the Noometry Index (30.4 vs 30.9), so choose on price, context window or the category you care about most.

Last verified . 24 shared benchmarks.

Amazon Nova Micro Amazon

30.4

Rank #294 Confirmed

Llama 4 Maverick Meta

30.9

Rank #282 Confirmed

Summary

  • They share 24 benchmarks with published results for both. Amazon Nova Micro scores higher in 5 categories and Llama 4 Maverick in 4 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in instruction following, where Llama 4 Maverick leads 71.7 to 56.3.
  • The biggest single-benchmark swing is MMLU-Pro: 51.1% for Amazon Nova Micro and 81% for Llama 4 Maverick.
  • Amazon Nova Micro is cheaper at $0.035 / $0.14 per million input/output tokens, against $0.19 / $0.65 for Llama 4 Maverick.
  • Llama 4 Maverick has downloadable open weights; the other is API-only.

Side by side

Amazon Nova Micro and Llama 4 Maverick specifications
Amazon Nova MicroLlama 4 Maverick
ProviderAmazonMeta
Noometry Index30.430.9
Released2024-12-032025-04-05
WeightsProprietaryOpen
Context window128K128K
Max output10K4K
Input $ / M tokens$0.035$0.19
Output $ / M tokens$0.14$0.65
Results tracked3254

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Amazon Nova Micro leads

Amazon Nova Micro: 30.5 (#295), Llama 4 Maverick: 26.6 (#324)

Coding benchmarks
BenchmarkAmazon Nova MicroLlama 4 Maverick
LMArena Coding12181302
SWE-bench Verified (bash only)—21%
Aider Polyglot—15.6%
SciCode—33.1%
WeirdML—24.5%
BigCodeBench Instruct—49.7%
LiveBench Coding20.2%—
BigCodeBench Complete—61.4%
ALE-Bench—172.97

Agentic & Tool Use Llama 4 Maverick leads

Amazon Nova Micro: 22.1 (#132), Llama 4 Maverick: 28.2 (#91)

Agentic & Tool Use benchmarks
BenchmarkAmazon Nova MicroLlama 4 Maverick
Berkeley Function Calling Leaderboard22.3%37.3%

Reasoning Amazon Nova Micro leads

Amazon Nova Micro: 17.4 (#294), Llama 4 Maverick: 10.1 (#342)

Reasoning benchmarks
BenchmarkAmazon Nova MicroLlama 4 Maverick
LMArena Hard Prompts11911281
ARC-AGI-2—0%
SimpleBench—27.7%
Kagi LLM Benchmark—55.9%
NYT Connections (extended)—8%
ARC-AGI-1—4.4%
CritPt—0%
EnigmaEval—0.6%
LiveBench Reasoning25.1%—
DTBench—61.9%
LiveBench Data Analysis34%—
LMCA—15.9%
Epoch Capabilities Index—132.2
ForecastBench—57.5
LiveBench29.6%—

Math Too close to call

Amazon Nova Micro: 26.9 (#254), Llama 4 Maverick: 26.0 (#262)

Math benchmarks
BenchmarkAmazon Nova MicroLlama 4 Maverick
Omni-MATH21.4%42.2%
LMArena Math12061299
OTIS Mock AIME 2024-2025—20.6%
LiveBench Math34.5%—
MATH Level 5—73%
FrontierMath (Feb 2025 set)—0.7%

Knowledge Llama 4 Maverick leads

Amazon Nova Micro: 29.6 (#237), Llama 4 Maverick: 33.4 (#204)

Knowledge benchmarks
BenchmarkAmazon Nova MicroLlama 4 Maverick
MMLU-Pro51.1%81%
Vectara Hallucination Rate5.5%8.2%
GPQA (HELM)38.3%65%
LMArena Expert11841259
GPQA Diamond—67%
Humanity's Last Exam—5.7%
Confabulations—22.6%
MMLU70.8%—

Multimodal Not comparable

Amazon Nova Micro: —, Llama 4 Maverick: 31.6 (#105)

Multimodal benchmarks
BenchmarkAmazon Nova MicroLlama 4 Maverick
LMArena Vision—1142
GeoBench—52%
SpatialViz-Bench—31.8%

Multilingual Llama 4 Maverick leads

Amazon Nova Micro: 36.5 (#239), Llama 4 Maverick: 42.2 (#195)

Multilingual benchmarks
BenchmarkAmazon Nova MicroLlama 4 Maverick
LMArena Non-English11861269
LMArena Chinese12091277
LMArena French12381259
LMArena German11921291
LMArena Japanese11541207
LMArena Korean11501203
LMArena Russian11851286
LMArena Spanish12251293

Instruction Following Llama 4 Maverick leads

Amazon Nova Micro: 56.3 (#272), Llama 4 Maverick: 71.7 (#146)

Instruction Following benchmarks
BenchmarkAmazon Nova MicroLlama 4 Maverick
IFEval76%90.8%
LMArena Instruction Following11741267
LiveBench Instruction Following48%—

Long Context Amazon Nova Micro leads

Amazon Nova Micro: 36.5 (#229), Llama 4 Maverick: 31.4 (#279)

Long Context benchmarks
BenchmarkAmazon Nova MicroLlama 4 Maverick
LMArena Longer Query12051280
Fiction.LiveBench—46.2%

Writing & Preference Too close to call

Amazon Nova Micro: 39.5 (#247), Llama 4 Maverick: 38.8 (#252)

Writing & Preference benchmarks
BenchmarkAmazon Nova MicroLlama 4 Maverick
LMArena Text12081287
LMArena Creative Writing11721267
WildBench74.3%80%
LMArena Multi-Turn11781289
Short-Story Creative Writing—62%
EQ-Bench Creative Writing—860
LiveBench Language15.8%—

Frequently asked questions

Is Amazon Nova Micro better than Llama 4 Maverick?

Amazon Nova Micro and Llama 4 Maverick score almost the same on the Noometry Index (30.4 vs 30.9), so choose on price, context window or the category you care about most.

Which is cheaper, Amazon Nova Micro or Llama 4 Maverick?

Amazon Nova Micro is cheaper. It lists at $0.035 per million input tokens and $0.14 per million output tokens; Llama 4 Maverick lists at $0.19 and $0.65.

Is Amazon Nova Micro or Llama 4 Maverick better for coding?

Amazon Nova Micro scores higher on coding benchmarks: 30.5 versus 26.6 in the Noometry coding category.

Which has the bigger context window?

Both accept 128K tokens.

How many benchmarks do Amazon Nova Micro and Llama 4 Maverick share?

24 benchmarks have published results for both models. Amazon Nova Micro has 32 scored results on Noometry and Llama 4 Maverick has 54.

Related comparisons

Go deeper