Model comparison

Command R vs Llama2 70b Steerlm Chat

Command R and Llama2 70b Steerlm Chat score almost the same on the Noometry Index (31.4 vs 31.8), so choose on price, context window or the category you care about most.

Last verified . 9 shared benchmarks.

Command R Cohere

31.4

Rank #272 Confirmed

Llama2 70b Steerlm Chat NVIDIA

31.8

Rank #268 Confirmed

Summary

  • They share 9 benchmarks with published results for both. Command R scores higher in 4 categories and Llama2 70b Steerlm Chat in 3 categories; 6 gaps are clear of the uncertainty.
  • The widest gap is in multilingual, where Command R leads 35.7 to 28.8.

Side by side

Command R and Llama2 70b Steerlm Chat specifications
Command RLlama2 70b Steerlm Chat
ProviderCohereNVIDIA
Noometry Index31.431.8
Released2024-08-30—
WeightsOpenOpen
Context window128K—
Max output4K—
Input $ / M tokens$0.15—
Output $ / M tokens$0.60—
Results tracked299

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Too close to call

Command R: 29.3 (#306), Llama2 70b Steerlm Chat: 29.9 (#300)

Coding benchmarks
BenchmarkCommand RLlama2 70b Steerlm Chat
LMArena Coding11691025
BigCodeBench Instruct37.1%—
LiveBench Coding17.9%—
BigCodeBench Complete45.2%—

Reasoning Llama2 70b Steerlm Chat leads

Command R: 13.8 (#331), Llama2 70b Steerlm Chat: 20.0 (#246)

Reasoning benchmarks
BenchmarkCommand RLlama2 70b Steerlm Chat
LMArena Hard Prompts11641047
LiveBench Reasoning21.9%—
DTBench46.4%—
LiveBench Data Analysis33.3%—
LMCA9.2%—
LiveBench27.5%—

Math Llama2 70b Steerlm Chat leads

Command R: 28.0 (#246), Llama2 70b Steerlm Chat: 31.3 (#226)

Math benchmarks
BenchmarkCommand RLlama2 70b Steerlm Chat
LMArena Math11551072
LiveBench Math19.4%—

Knowledge Not comparable

Command R: 31.0 (#221), Llama2 70b Steerlm Chat: —

Knowledge benchmarks
BenchmarkCommand RLlama2 70b Steerlm Chat
LMArena Expert1138—
MMLU65.2%—

Multilingual Command R leads

Command R: 35.7 (#245), Llama2 70b Steerlm Chat: 28.8 (#270)

Multilingual benchmarks
BenchmarkCommand RLlama2 70b Steerlm Chat
LMArena Non-English11741063
LMArena Chinese1182—
LMArena French1162—
LMArena German1176—
LMArena Japanese1143—
LMArena Korean1163—
LMArena Russian1174—
LMArena Spanish1151—

Instruction Following Command R leads

Command R: 58.1 (#261), Llama2 70b Steerlm Chat: 54.2 (#279)

Instruction Following benchmarks
BenchmarkCommand RLlama2 70b Steerlm Chat
LMArena Instruction Following11671060
LiveBench Instruction Following55.6%—

Long Context Command R leads

Command R: 36.3 (#231), Llama2 70b Steerlm Chat: 30.4 (#288)

Long Context benchmarks
BenchmarkCommand RLlama2 70b Steerlm Chat
LMArena Longer Query1198998

Writing & Preference Command R leads

Command R: 38.2 (#254), Llama2 70b Steerlm Chat: 31.6 (#283)

Writing & Preference benchmarks
BenchmarkCommand RLlama2 70b Steerlm Chat
LMArena Text11871098
LMArena Creative Writing11701091
LMArena Multi-Turn11631058
LiveBench Language16.7%—

Frequently asked questions

Is Command R better than Llama2 70b Steerlm Chat?

Command R and Llama2 70b Steerlm Chat score almost the same on the Noometry Index (31.4 vs 31.8), so choose on price, context window or the category you care about most.

Is Command R or Llama2 70b Steerlm Chat better for coding?

They score almost the same on coding (29.3 vs 29.9); test both on your own repository before choosing.

How many benchmarks do Command R and Llama2 70b Steerlm Chat share?

9 benchmarks have published results for both models. Command R has 29 scored results on Noometry and Llama2 70b Steerlm Chat has 9.

Related comparisons

Go deeper