Model comparison

DeepSeek-R1-Distill-Llama-70B vs Mercury

DeepSeek-R1-Distill-Llama-70B and Mercury score almost the same on the Noometry Index (37.8 vs 37.6), so choose on price, context window or the category you care about most.

Last verified . 1 shared benchmarks.

Mercury Inception

37.6

Rank #199 Confirmed

Summary

  • They share 1 benchmark with published results for both. DeepSeek-R1-Distill-Llama-70B scores higher in 3 categories and Mercury in 1 category; 4 gaps are clear of the uncertainty.
  • The widest gap is in reasoning, where DeepSeek-R1-Distill-Llama-70B leads 24.9 to 17.5.
  • The biggest single-benchmark swing is Kagi LLM Benchmark: 52.3% for DeepSeek-R1-Distill-Llama-70B and 21.6% for Mercury.
  • DeepSeek-R1-Distill-Llama-70B has downloadable open weights; the other is API-only.

Side by side

DeepSeek-R1-Distill-Llama-70B and Mercury specifications
DeepSeek-R1-Distill-Llama-70BMercury
ProviderDeepSeekInception
Noometry Index37.837.6
Released2025-01-20—
WeightsOpenProprietary
Context window——
Max output——
Input $ / M tokens——
Output $ / M tokens——
Results tracked139

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Mercury leads

DeepSeek-R1-Distill-Llama-70B: 36.8 (#202), Mercury: 38.7 (#170)

Coding benchmarks
BenchmarkDeepSeek-R1-Distill-Llama-70BMercury
BigCodeBench Instruct35.3%—
LiveBench Coding51.6%—
LMArena Coding—1322
BigCodeBench Complete49.9%—

Reasoning DeepSeek-R1-Distill-Llama-70B leads

DeepSeek-R1-Distill-Llama-70B: 24.9 (#156), Mercury: 17.5 (#293)

Reasoning benchmarks
BenchmarkDeepSeek-R1-Distill-Llama-70BMercury
Kagi LLM Benchmark52.3%21.6%
LiveBench Reasoning67.6%—
LMArena Hard Prompts—1285
LiveBench Data Analysis55.9%—
LiveBench54.5%—

Math Not comparable

DeepSeek-R1-Distill-Llama-70B: 36.0 (#176), Mercury: —

Math benchmarks
BenchmarkDeepSeek-R1-Distill-Llama-70BMercury
OTIS Mock AIME 2024-202551.4%—
LiveBench Math58.1%—
MATH Level 589.9%—

Knowledge Not comparable

DeepSeek-R1-Distill-Llama-70B: 30.7 (#225), Mercury: —

Knowledge benchmarks
BenchmarkDeepSeek-R1-Distill-Llama-70BMercury
GPQA Diamond55.7%—

Multilingual Not comparable

DeepSeek-R1-Distill-Llama-70B: —, Mercury: 41.6 (#206)

Multilingual benchmarks
BenchmarkDeepSeek-R1-Distill-Llama-70BMercury
LMArena Non-English—1260

Instruction Following DeepSeek-R1-Distill-Llama-70B leads

DeepSeek-R1-Distill-Llama-70B: 68.2 (#190), Mercury: 65.2 (#224)

Instruction Following benchmarks
BenchmarkDeepSeek-R1-Distill-Llama-70BMercury
LiveBench Instruction Following69.9%—
LMArena Instruction Following—1239

Long Context Not comparable

DeepSeek-R1-Distill-Llama-70B: —, Mercury: 38.4 (#198)

Long Context benchmarks
BenchmarkDeepSeek-R1-Distill-Llama-70BMercury
LMArena Longer Query—1266

Writing & Preference DeepSeek-R1-Distill-Llama-70B leads

DeepSeek-R1-Distill-Llama-70B: 49.0 (#194), Mercury: 46.2 (#221)

Writing & Preference benchmarks
BenchmarkDeepSeek-R1-Distill-Llama-70BMercury
LMArena Text—1282
LMArena Creative Writing—1191
LMArena Multi-Turn—1282
LiveBench Language23.8%—

Frequently asked questions

Is DeepSeek-R1-Distill-Llama-70B better than Mercury?

DeepSeek-R1-Distill-Llama-70B and Mercury score almost the same on the Noometry Index (37.8 vs 37.6), so choose on price, context window or the category you care about most.

Is DeepSeek-R1-Distill-Llama-70B or Mercury better for coding?

Mercury scores higher on coding benchmarks: 38.7 versus 36.8 in the Noometry coding category.

How many benchmarks do DeepSeek-R1-Distill-Llama-70B and Mercury share?

1 benchmark has published results for both models. DeepSeek-R1-Distill-Llama-70B has 13 scored results on Noometry and Mercury has 9.

Related comparisons

Go deeper