Model comparison

Command A vs DeepSeek-V3.1-Terminus

DeepSeek-V3.1-Terminus is the stronger model overall, scoring 43.1 to 36.5 on the Noometry Index.

Last verified . 13 shared benchmarks.

Command A Cohere

36.5

Rank #215 Confirmed

DeepSeek-V3.1-Terminus DeepSeek

43.1

Rank #97 Confirmed

Summary

  • They share 13 benchmarks with published results for both. Command A scores higher in 0 categories and DeepSeek-V3.1-Terminus in 7 categories; 7 gaps are clear of the uncertainty.
  • The widest gap is in coding, where DeepSeek-V3.1-Terminus leads 42.0 to 27.2.
  • The biggest single-benchmark swing is Kagi LLM Benchmark: 28.8% for Command A and 57.4% for DeepSeek-V3.1-Terminus.
  • DeepSeek-V3.1-Terminus is cheaper at $0.27 / $1 per million input/output tokens, against $2.50 / $10 for Command A.
  • Command A accepts more context: 256K tokens versus 164K.

Side by side

Command A and DeepSeek-V3.1-Terminus specifications
Command ADeepSeek-V3.1-Terminus
ProviderCohereDeepSeek
Noometry Index36.543.1
Released2025-03-132025-09-22
WeightsOpenOpen
Context window256K164K
Max output8K147K
Input $ / M tokens$2.50$0.27
Output $ / M tokens$10$1
Results tracked2416

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding DeepSeek-V3.1-Terminus leads

Command A: 27.2 (#322), DeepSeek-V3.1-Terminus: 42.0 (#113)

Coding benchmarks
BenchmarkCommand ADeepSeek-V3.1-Terminus
LMArena Coding13301426
Aider Polyglot12%—
SciCode—40.6%
ALE-Bench—745.17

Agentic & Tool Use Not comparable

Command A: 35.9 (#40), DeepSeek-V3.1-Terminus: —

Agentic & Tool Use benchmarks
BenchmarkCommand ADeepSeek-V3.1-Terminus
Berkeley Function Calling Leaderboard57.1%—

Reasoning DeepSeek-V3.1-Terminus leads

Command A: 18.3 (#283), DeepSeek-V3.1-Terminus: 26.4 (#133)

Reasoning benchmarks
BenchmarkCommand ADeepSeek-V3.1-Terminus
Kagi LLM Benchmark28.8%57.4%
LMArena Hard Prompts13261426
DTBench61.3%81.3%
LMCA10.3%28.6%
CritPt—1.7%

Math DeepSeek-V3.1-Terminus leads

Command A: 36.2 (#171), DeepSeek-V3.1-Terminus: 38.5 (#137)

Math benchmarks
BenchmarkCommand ADeepSeek-V3.1-Terminus
LMArena Math13001402

Knowledge Not comparable

Command A: 37.1 (#159), DeepSeek-V3.1-Terminus: —

Knowledge benchmarks
BenchmarkCommand ADeepSeek-V3.1-Terminus
Vectara Hallucination Rate9.3%—
LMArena Expert1295—

Multilingual DeepSeek-V3.1-Terminus leads

Command A: 45.3 (#170), DeepSeek-V3.1-Terminus: 52.1 (#92)

Multilingual benchmarks
BenchmarkCommand ADeepSeek-V3.1-Terminus
LMArena Non-English13131407
LMArena Russian13141436
LMArena Chinese1327—
LMArena French1351—
LMArena German1341—
LMArena Japanese1285—
LMArena Korean1285—
LMArena Spanish1347—

Instruction Following DeepSeek-V3.1-Terminus leads

Command A: 69.1 (#177), DeepSeek-V3.1-Terminus: 74.0 (#106)

Instruction Following benchmarks
BenchmarkCommand ADeepSeek-V3.1-Terminus
LMArena Instruction Following13091404

Long Context DeepSeek-V3.1-Terminus leads

Command A: 40.6 (#151), DeepSeek-V3.1-Terminus: 43.4 (#97)

Long Context benchmarks
BenchmarkCommand ADeepSeek-V3.1-Terminus
LMArena Longer Query13341421

Writing & Preference DeepSeek-V3.1-Terminus leads

Command A: 47.6 (#208), DeepSeek-V3.1-Terminus: 61.0 (#92)

Writing & Preference benchmarks
BenchmarkCommand ADeepSeek-V3.1-Terminus
LMArena Text13311419
LMArena Creative Writing13191403
LMArena Multi-Turn13391411
EQ-Bench Creative Writing1145—

Frequently asked questions

Is Command A better than DeepSeek-V3.1-Terminus?

DeepSeek-V3.1-Terminus is the stronger model overall, scoring 43.1 to 36.5 on the Noometry Index.

Which is cheaper, Command A or DeepSeek-V3.1-Terminus?

DeepSeek-V3.1-Terminus is cheaper. It lists at $0.27 per million input tokens and $1 per million output tokens; Command A lists at $2.50 and $10.

Is Command A or DeepSeek-V3.1-Terminus better for coding?

DeepSeek-V3.1-Terminus scores higher on coding benchmarks: 42.0 versus 27.2 in the Noometry coding category.

Which has the bigger context window?

Command A does, with 256K tokens against 164K.

How many benchmarks do Command A and DeepSeek-V3.1-Terminus share?

13 benchmarks have published results for both models. Command A has 24 scored results on Noometry and DeepSeek-V3.1-Terminus has 16.

Related comparisons

Go deeper