Model comparison

Claude 2.1 vs Dolly 2.0-12b

Claude 2.1 and Dolly 2.0-12b score almost the same on the Noometry Index (25.2 vs 25.5), so choose on price, context window or the category you care about most.

Last verified . 2 shared benchmarks.

Claude 2.1 Anthropic

25.2

Rank #345 Reported

Dolly 2.0-12b Databricks

25.5

Rank #342 Confirmed

Summary

  • They share 2 benchmarks with published results for both. Claude 2.1 scores higher in 2 categories and Dolly 2.0-12b in 1 category; 3 gaps are clear of the uncertainty.
  • The widest gap is in math, where Dolly 2.0-12b leads 27.3 to 10.2.
  • Dolly 2.0-12b has downloadable open weights; the other is API-only.

Side by side

Claude 2.1 and Dolly 2.0-12b specifications
Claude 2.1Dolly 2.0-12b
ProviderAnthropicDatabricks
Noometry Index25.225.5
Released2023-11-212023-04-11
WeightsProprietaryOpen
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked717

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Claude 2.1 leads

Claude 2.1: 26.2 (#327), Dolly 2.0-12b: 23.4 (#332)

Coding benchmarks
BenchmarkClaude 2.1Dolly 2.0-12b
WeirdML7.1%—
LMArena Coding—776

Reasoning Claude 2.1 leads

Claude 2.1: 21.4 (#221), Dolly 2.0-12b: 15.3 (#316)

Reasoning benchmarks
BenchmarkClaude 2.1Dolly 2.0-12b
Epoch Capabilities Index119.2789.67
LMArena Hard Prompts—804
DTBench51%—
ForecastBench54.2—
HellaSwag—70.8%
PIQA—75.4%
WinoGrande—61.8%

Math Dolly 2.0-12b leads

Claude 2.1: 10.2 (#315), Dolly 2.0-12b: 27.3 (#251)

Math benchmarks
BenchmarkClaude 2.1Dolly 2.0-12b
OTIS Mock AIME 2024-20251.9%—
LMArena Math—871

Knowledge Not comparable

Claude 2.1: 15.4 (#292), Dolly 2.0-12b: —

Knowledge benchmarks
BenchmarkClaude 2.1Dolly 2.0-12b
MMLU73.5%26.2%
GPQA Diamond33%—
ARC (AI2) Challenge—39.6%
BoolQ—56.3%
OpenBookQA—39.2%

Multilingual Not comparable

Claude 2.1: —, Dolly 2.0-12b: 17.4 (#296)

Multilingual benchmarks
BenchmarkClaude 2.1Dolly 2.0-12b
LMArena Non-English—836
LMArena Chinese—836

Instruction Following Not comparable

Claude 2.1: —, Dolly 2.0-12b: 38.7 (#304)

Instruction Following benchmarks
BenchmarkClaude 2.1Dolly 2.0-12b
LMArena Instruction Following—814

Writing & Preference Not comparable

Claude 2.1: —, Dolly 2.0-12b: 15.2 (#311)

Writing & Preference benchmarks
BenchmarkClaude 2.1Dolly 2.0-12b
LMArena Text—851
LMArena Creative Writing—864
LMArena Multi-Turn—740

Frequently asked questions

Is Claude 2.1 better than Dolly 2.0-12b?

Claude 2.1 and Dolly 2.0-12b score almost the same on the Noometry Index (25.2 vs 25.5), so choose on price, context window or the category you care about most.

Is Claude 2.1 or Dolly 2.0-12b better for coding?

Claude 2.1 scores higher on coding benchmarks: 26.2 versus 23.4 in the Noometry coding category.

How many benchmarks do Claude 2.1 and Dolly 2.0-12b share?

2 benchmarks have published results for both models. Claude 2.1 has 7 scored results on Noometry and Dolly 2.0-12b has 17.

Related comparisons

Go deeper