Model comparison

Dolly 2.0-12b vs Mistral Nemo

Dolly 2.0-12b and Mistral Nemo score almost the same on the Noometry Index (25.5 vs 26.4), so choose on price, context window or the category you care about most.

Last verified . 3 shared benchmarks.

Dolly 2.0-12b Databricks

25.5

Rank #342 Confirmed

Mistral Nemo Mistral AI

26.4

Rank #337 Confirmed

Summary

  • They share 3 benchmarks with published results for both. Dolly 2.0-12b scores higher in 1 category and Mistral Nemo in 2 categories; 3 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Mistral Nemo leads 28.5 to 15.2.

Side by side

Dolly 2.0-12b and Mistral Nemo specifications
Dolly 2.0-12bMistral Nemo
ProviderDatabricksMistral AI
Noometry Index25.526.4
Released2023-04-112024-07-01
WeightsOpenOpen
Context window—128K
Max output—128K
Input $ / M tokens—$0.15
Output $ / M tokens—$0.15
Results tracked1710

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Not comparable

Dolly 2.0-12b: 23.4 (#332), Mistral Nemo: —

Coding benchmarks
BenchmarkDolly 2.0-12bMistral Nemo
LMArena Coding776—

Agentic & Tool Use Not comparable

Dolly 2.0-12b: —, Mistral Nemo: 23.5 (#125)

Agentic & Tool Use benchmarks
BenchmarkDolly 2.0-12bMistral Nemo
Berkeley Function Calling Leaderboard—27.6%
BALROG—17.6%

Reasoning Mistral Nemo leads

Dolly 2.0-12b: 15.3 (#316), Mistral Nemo: 20.7 (#232)

Reasoning benchmarks
BenchmarkDolly 2.0-12bMistral Nemo
Epoch Capabilities Index89.67118.68
PIQA75.4%83.5%
LMArena Hard Prompts804—
DTBench—48.6%
HellaSwag70.8%—
WinoGrande61.8%—

Math Dolly 2.0-12b leads

Dolly 2.0-12b: 27.3 (#251), Mistral Nemo: 25.5 (#268)

Math benchmarks
BenchmarkDolly 2.0-12bMistral Nemo
LMArena Math871—
MATH Level 5—10.8%
GSM8K—84.2%

Knowledge Not comparable

Dolly 2.0-12b: —, Mistral Nemo: 12.3 (#298)

Knowledge benchmarks
BenchmarkDolly 2.0-12bMistral Nemo
BoolQ56.3%82.5%
GPQA Diamond—29.9%
ARC (AI2) Challenge39.6%—
MMLU26.2%—
OpenBookQA39.2%—

Multilingual Not comparable

Dolly 2.0-12b: 17.4 (#296), Mistral Nemo: —

Multilingual benchmarks
BenchmarkDolly 2.0-12bMistral Nemo
LMArena Non-English836—
LMArena Chinese836—

Instruction Following Not comparable

Dolly 2.0-12b: 38.7 (#304), Mistral Nemo: —

Instruction Following benchmarks
BenchmarkDolly 2.0-12bMistral Nemo
LMArena Instruction Following814—

Writing & Preference Mistral Nemo leads

Dolly 2.0-12b: 15.2 (#311), Mistral Nemo: 28.5 (#296)

Writing & Preference benchmarks
BenchmarkDolly 2.0-12bMistral Nemo
LMArena Text851—
LMArena Creative Writing864—
EQ-Bench Creative Writing—881
LMArena Multi-Turn740—

Frequently asked questions

Is Dolly 2.0-12b better than Mistral Nemo?

Dolly 2.0-12b and Mistral Nemo score almost the same on the Noometry Index (25.5 vs 26.4), so choose on price, context window or the category you care about most.

How many benchmarks do Dolly 2.0-12b and Mistral Nemo share?

3 benchmarks have published results for both models. Dolly 2.0-12b has 17 scored results on Noometry and Mistral Nemo has 10.

Related comparisons

Go deeper