Model comparison
Command A vs Llama 3.1 Nemotron Ultra 253b v1
Command A and Llama 3.1 Nemotron Ultra 253b v1 score almost the same on the Noometry Index (36.5 vs 36.7), so choose on price, context window or the category you care about most.
Last verified . 11 shared benchmarks.
Summary
- They share 11 benchmarks with published results for both. Command A scores higher in 4 categories and Llama 3.1 Nemotron Ultra 253b v1 in 4 categories; 7 gaps are clear of the uncertainty.
- The widest gap is in agentic & tool use, where Command A leads 35.9 to 15.7.
- The biggest single-benchmark swing is Berkeley Function Calling Leaderboard: 57.1% for Command A and 10% for Llama 3.1 Nemotron Ultra 253b v1.
Side by side
| Command A | Llama 3.1 Nemotron Ultra 253b v1 | |
|---|---|---|
| Provider | Cohere | NVIDIA |
| Noometry Index | 36.5 | 36.7 |
| Released | 2025-03-13 | — |
| Weights | Open | Open |
| Context window | 256K | — |
| Max output | 8K | — |
| Input $ / M tokens | $2.50 | — |
| Output $ / M tokens | $10 | — |
| Results tracked | 24 | 11 |
Sponsored placements are available on pages like this one. Advertise on Noometry
Category by category
Coding Llama 3.1 Nemotron Ultra 253b v1 leads
Command A: 27.2 (#322), Llama 3.1 Nemotron Ultra 253b v1: 38.4 (#177)
| Benchmark | Command A | Llama 3.1 Nemotron Ultra 253b v1 |
|---|---|---|
| LMArena Coding | 1330 | 1312 |
| Aider Polyglot | 12% | — |
Agentic & Tool Use Command A leads
Command A: 35.9 (#40), Llama 3.1 Nemotron Ultra 253b v1: 15.7 (#149)
| Benchmark | Command A | Llama 3.1 Nemotron Ultra 253b v1 |
|---|---|---|
| Berkeley Function Calling Leaderboard | 57.1% | 10% |
Reasoning Llama 3.1 Nemotron Ultra 253b v1 leads
Command A: 18.3 (#283), Llama 3.1 Nemotron Ultra 253b v1: 26.3 (#134)
| Benchmark | Command A | Llama 3.1 Nemotron Ultra 253b v1 |
|---|---|---|
| LMArena Hard Prompts | 1326 | 1316 |
| Kagi LLM Benchmark | 28.8% | — |
| DTBench | 61.3% | — |
| LMCA | 10.3% | — |
Math Llama 3.1 Nemotron Ultra 253b v1 leads
Command A: 36.2 (#171), Llama 3.1 Nemotron Ultra 253b v1: 37.5 (#152)
| Benchmark | Command A | Llama 3.1 Nemotron Ultra 253b v1 |
|---|---|---|
| LMArena Math | 1300 | 1360 |
Knowledge Not comparable
Command A: 37.1 (#159), Llama 3.1 Nemotron Ultra 253b v1: —
| Benchmark | Command A | Llama 3.1 Nemotron Ultra 253b v1 |
|---|---|---|
| Vectara Hallucination Rate | 9.3% | — |
| LMArena Expert | 1295 | — |
Multilingual Command A leads
Command A: 45.3 (#170), Llama 3.1 Nemotron Ultra 253b v1: 43.1 (#187)
| Benchmark | Command A | Llama 3.1 Nemotron Ultra 253b v1 |
|---|---|---|
| LMArena Non-English | 1313 | 1282 |
| LMArena Russian | 1314 | 1284 |
| LMArena Chinese | 1327 | — |
| LMArena French | 1351 | — |
| LMArena German | 1341 | — |
| LMArena Japanese | 1285 | — |
| LMArena Korean | 1285 | — |
| LMArena Spanish | 1347 | — |
Instruction Following Too close to call
Command A: 69.1 (#177), Llama 3.1 Nemotron Ultra 253b v1: 69.0 (#178)
| Benchmark | Command A | Llama 3.1 Nemotron Ultra 253b v1 |
|---|---|---|
| LMArena Instruction Following | 1309 | 1308 |
Long Context Command A leads
Command A: 40.6 (#151), Llama 3.1 Nemotron Ultra 253b v1: 39.5 (#177)
| Benchmark | Command A | Llama 3.1 Nemotron Ultra 253b v1 |
|---|---|---|
| LMArena Longer Query | 1334 | 1299 |
Writing & Preference Llama 3.1 Nemotron Ultra 253b v1 leads
Command A: 47.6 (#208), Llama 3.1 Nemotron Ultra 253b v1: 52.2 (#175)
| Benchmark | Command A | Llama 3.1 Nemotron Ultra 253b v1 |
|---|---|---|
| LMArena Text | 1331 | 1320 |
| LMArena Creative Writing | 1319 | 1314 |
| LMArena Multi-Turn | 1339 | 1317 |
| EQ-Bench Creative Writing | 1145 | — |
Frequently asked questions
Is Command A better than Llama 3.1 Nemotron Ultra 253b v1?
Command A and Llama 3.1 Nemotron Ultra 253b v1 score almost the same on the Noometry Index (36.5 vs 36.7), so choose on price, context window or the category you care about most.
Is Command A or Llama 3.1 Nemotron Ultra 253b v1 better for coding?
Llama 3.1 Nemotron Ultra 253b v1 scores higher on coding benchmarks: 38.4 versus 27.2 in the Noometry coding category.
How many benchmarks do Command A and Llama 3.1 Nemotron Ultra 253b v1 share?
11 benchmarks have published results for both models. Command A has 24 scored results on Noometry and Llama 3.1 Nemotron Ultra 253b v1 has 11.