Model comparison

Command R vs DeepSeek-V3.1-Terminus

DeepSeek-V3.1-Terminus is the stronger model overall, scoring 43.1 to 31.4 on the Noometry Index. Command R costs 1.7× less per token, which makes it the better buy when DeepSeek-V3.1-Terminus's lead doesn't matter for your workload.

Last verified . 12 shared benchmarks.

Command R Cohere

31.4

Rank #272 Confirmed

DeepSeek-V3.1-Terminus DeepSeek

43.1

Rank #97 Confirmed

Summary

  • They share 12 benchmarks with published results for both. Command R scores higher in 0 categories and DeepSeek-V3.1-Terminus in 7 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where DeepSeek-V3.1-Terminus leads 61.0 to 38.2.
  • The biggest single-benchmark swing is DTBench: 46.4% for Command R and 81.3% for DeepSeek-V3.1-Terminus.
  • Command R is cheaper at $0.15 / $0.60 per million input/output tokens, against $0.27 / $1 for DeepSeek-V3.1-Terminus.
  • DeepSeek-V3.1-Terminus accepts more context: 164K tokens versus 128K.

Side by side

Command R and DeepSeek-V3.1-Terminus specifications
Command RDeepSeek-V3.1-Terminus
ProviderCohereDeepSeek
Noometry Index31.443.1
Released2024-08-302025-09-22
WeightsOpenOpen
Context window128K164K
Max output4K147K
Input $ / M tokens$0.15$0.27
Output $ / M tokens$0.60$1
Results tracked2916

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding DeepSeek-V3.1-Terminus leads

Command R: 29.3 (#306), DeepSeek-V3.1-Terminus: 42.0 (#113)

Coding benchmarks
BenchmarkCommand RDeepSeek-V3.1-Terminus
LMArena Coding11691426
SciCode—40.6%
BigCodeBench Instruct37.1%—
LiveBench Coding17.9%—
BigCodeBench Complete45.2%—
ALE-Bench—745.17

Reasoning DeepSeek-V3.1-Terminus leads

Command R: 13.8 (#331), DeepSeek-V3.1-Terminus: 26.4 (#133)

Reasoning benchmarks
BenchmarkCommand RDeepSeek-V3.1-Terminus
LMArena Hard Prompts11641426
DTBench46.4%81.3%
LMCA9.2%28.6%
Kagi LLM Benchmark—57.4%
CritPt—1.7%
LiveBench Reasoning21.9%—
LiveBench Data Analysis33.3%—
LiveBench27.5%—

Math DeepSeek-V3.1-Terminus leads

Command R: 28.0 (#246), DeepSeek-V3.1-Terminus: 38.5 (#137)

Math benchmarks
BenchmarkCommand RDeepSeek-V3.1-Terminus
LMArena Math11551402
LiveBench Math19.4%—

Knowledge Not comparable

Command R: 31.0 (#221), DeepSeek-V3.1-Terminus: —

Knowledge benchmarks
BenchmarkCommand RDeepSeek-V3.1-Terminus
LMArena Expert1138—
MMLU65.2%—

Multilingual DeepSeek-V3.1-Terminus leads

Command R: 35.7 (#245), DeepSeek-V3.1-Terminus: 52.1 (#92)

Multilingual benchmarks
BenchmarkCommand RDeepSeek-V3.1-Terminus
LMArena Non-English11741407
LMArena Russian11741436
LMArena Chinese1182—
LMArena French1162—
LMArena German1176—
LMArena Japanese1143—
LMArena Korean1163—
LMArena Spanish1151—

Instruction Following DeepSeek-V3.1-Terminus leads

Command R: 58.1 (#261), DeepSeek-V3.1-Terminus: 74.0 (#106)

Instruction Following benchmarks
BenchmarkCommand RDeepSeek-V3.1-Terminus
LMArena Instruction Following11671404
LiveBench Instruction Following55.6%—

Long Context DeepSeek-V3.1-Terminus leads

Command R: 36.3 (#231), DeepSeek-V3.1-Terminus: 43.4 (#97)

Long Context benchmarks
BenchmarkCommand RDeepSeek-V3.1-Terminus
LMArena Longer Query11981421

Writing & Preference DeepSeek-V3.1-Terminus leads

Command R: 38.2 (#254), DeepSeek-V3.1-Terminus: 61.0 (#92)

Writing & Preference benchmarks
BenchmarkCommand RDeepSeek-V3.1-Terminus
LMArena Text11871419
LMArena Creative Writing11701403
LMArena Multi-Turn11631411
LiveBench Language16.7%—

Frequently asked questions

Is Command R better than DeepSeek-V3.1-Terminus?

DeepSeek-V3.1-Terminus is the stronger model overall, scoring 43.1 to 31.4 on the Noometry Index. Command R costs 1.7× less per token, which makes it the better buy when DeepSeek-V3.1-Terminus's lead doesn't matter for your workload.

Which is cheaper, Command R or DeepSeek-V3.1-Terminus?

Command R is cheaper. It lists at $0.15 per million input tokens and $0.60 per million output tokens; DeepSeek-V3.1-Terminus lists at $0.27 and $1.

Is Command R or DeepSeek-V3.1-Terminus better for coding?

DeepSeek-V3.1-Terminus scores higher on coding benchmarks: 42.0 versus 29.3 in the Noometry coding category.

Which has the bigger context window?

DeepSeek-V3.1-Terminus does, with 164K tokens against 128K.

How many benchmarks do Command R and DeepSeek-V3.1-Terminus share?

12 benchmarks have published results for both models. Command R has 29 scored results on Noometry and DeepSeek-V3.1-Terminus has 16.

Related comparisons

Go deeper