Model comparison

Amazon Nova Micro vs Dolly 2.0-12b

Amazon Nova Micro is the stronger model overall, scoring 30.4 to 25.5 on the Noometry Index.

Last verified . 10 shared benchmarks.

Amazon Nova Micro Amazon

30.4

Rank #294 Confirmed

Dolly 2.0-12b Databricks

25.5

Rank #342 Confirmed

Summary

  • They share 10 benchmarks with published results for both. Amazon Nova Micro scores higher in 5 categories and Dolly 2.0-12b in 1 category; 5 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Amazon Nova Micro leads 39.5 to 15.2.
  • Dolly 2.0-12b has downloadable open weights; the other is API-only.

Side by side

Amazon Nova Micro and Dolly 2.0-12b specifications
Amazon Nova MicroDolly 2.0-12b
ProviderAmazonDatabricks
Noometry Index30.425.5
Released2024-12-032023-04-11
WeightsProprietaryOpen
Context window128K—
Max output10K—
Input $ / M tokens$0.035—
Output $ / M tokens$0.14—
Results tracked3217

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Amazon Nova Micro leads

Amazon Nova Micro: 30.5 (#295), Dolly 2.0-12b: 23.4 (#332)

Coding benchmarks
BenchmarkAmazon Nova MicroDolly 2.0-12b
LMArena Coding1218776
LiveBench Coding20.2%—

Agentic & Tool Use Not comparable

Amazon Nova Micro: 22.1 (#132), Dolly 2.0-12b: —

Agentic & Tool Use benchmarks
BenchmarkAmazon Nova MicroDolly 2.0-12b
Berkeley Function Calling Leaderboard22.3%—

Reasoning Amazon Nova Micro leads

Amazon Nova Micro: 17.4 (#294), Dolly 2.0-12b: 15.3 (#316)

Reasoning benchmarks
BenchmarkAmazon Nova MicroDolly 2.0-12b
LMArena Hard Prompts1191804
LiveBench Reasoning25.1%—
LiveBench Data Analysis34%—
Epoch Capabilities Index—89.67
HellaSwag—70.8%
LiveBench29.6%—
PIQA—75.4%
WinoGrande—61.8%

Math Too close to call

Amazon Nova Micro: 26.9 (#254), Dolly 2.0-12b: 27.3 (#251)

Math benchmarks
BenchmarkAmazon Nova MicroDolly 2.0-12b
LMArena Math1206871
Omni-MATH21.4%—
LiveBench Math34.5%—

Knowledge Not comparable

Amazon Nova Micro: 29.6 (#237), Dolly 2.0-12b: —

Knowledge benchmarks
BenchmarkAmazon Nova MicroDolly 2.0-12b
MMLU70.8%26.2%
MMLU-Pro51.1%—
Vectara Hallucination Rate5.5%—
GPQA (HELM)38.3%—
LMArena Expert1184—
ARC (AI2) Challenge—39.6%
BoolQ—56.3%
OpenBookQA—39.2%

Multilingual Amazon Nova Micro leads

Amazon Nova Micro: 36.5 (#239), Dolly 2.0-12b: 17.4 (#296)

Multilingual benchmarks
BenchmarkAmazon Nova MicroDolly 2.0-12b
LMArena Non-English1186836
LMArena Chinese1209836
LMArena French1238—
LMArena German1192—
LMArena Japanese1154—
LMArena Korean1150—
LMArena Russian1185—
LMArena Spanish1225—

Instruction Following Amazon Nova Micro leads

Amazon Nova Micro: 56.3 (#272), Dolly 2.0-12b: 38.7 (#304)

Instruction Following benchmarks
BenchmarkAmazon Nova MicroDolly 2.0-12b
LMArena Instruction Following1174814
LiveBench Instruction Following48%—
IFEval76%—

Long Context Not comparable

Amazon Nova Micro: 36.5 (#229), Dolly 2.0-12b: —

Long Context benchmarks
BenchmarkAmazon Nova MicroDolly 2.0-12b
LMArena Longer Query1205—

Writing & Preference Amazon Nova Micro leads

Amazon Nova Micro: 39.5 (#247), Dolly 2.0-12b: 15.2 (#311)

Writing & Preference benchmarks
BenchmarkAmazon Nova MicroDolly 2.0-12b
LMArena Text1208851
LMArena Creative Writing1172864
LMArena Multi-Turn1178740
WildBench74.3%—
LiveBench Language15.8%—

Frequently asked questions

Is Amazon Nova Micro better than Dolly 2.0-12b?

Amazon Nova Micro is the stronger model overall, scoring 30.4 to 25.5 on the Noometry Index.

Is Amazon Nova Micro or Dolly 2.0-12b better for coding?

Amazon Nova Micro scores higher on coding benchmarks: 30.5 versus 23.4 in the Noometry coding category.

How many benchmarks do Amazon Nova Micro and Dolly 2.0-12b share?

10 benchmarks have published results for both models. Amazon Nova Micro has 32 scored results on Noometry and Dolly 2.0-12b has 17.

Related comparisons

Go deeper