Model comparison
Deepseek Coder v2 vs Mistral Medium 3.5
Mistral Medium 3.5 is the stronger model overall, scoring 40.2 to 35.9 on the Noometry Index.
Last verified . 16 shared benchmarks.
Summary
- They share 16 benchmarks with published results for both. Deepseek Coder v2 scores higher in 2 categories and Mistral Medium 3.5 in 6 categories; 8 gaps are clear of the uncertainty.
- The widest gap is in writing & preference, where Mistral Medium 3.5 leads 58.5 to 38.2.
Side by side
| Deepseek Coder v2 | Mistral Medium 3.5 | |
|---|---|---|
| Provider | DeepSeek | Mistral AI |
| Noometry Index | 35.9 | 40.2 |
| Released | 2024-06-17 | — |
| Weights | Open | Open |
| Context window | — | 262K |
| Max output | — | 210K |
| Input $ / M tokens | — | $1.50 |
| Output $ / M tokens | — | $7.50 |
| Results tracked | 24 | 22 |
Sponsored placements are available on pages like this one. Advertise on Noometry
Category by category
Coding Deepseek Coder v2 leads
Deepseek Coder v2: 38.1 (#183), Mistral Medium 3.5: 36.0 (#213)
| Benchmark | Deepseek Coder v2 | Mistral Medium 3.5 |
|---|---|---|
| LMArena Coding | 1251 | 1461 |
| LMArena WebDev | — | 1264 |
| BigCodeBench Instruct | 48.2% | — |
| BigCodeBench Complete | 59.7% | — |
| HumanEval+ | 82.3% | — |
| MBPP+ | 75.1% | — |
Reasoning Deepseek Coder v2 leads
Deepseek Coder v2: 23.6 (#176), Mistral Medium 3.5: 17.3 (#295)
| Benchmark | Deepseek Coder v2 | Mistral Medium 3.5 |
|---|---|---|
| LMArena Hard Prompts | 1207 | 1436 |
| Kagi LLM Benchmark | — | 41.4% |
| NYT Connections (extended) | — | 12.9% |
| Epoch Capabilities Index | — | 141.35 |
| WinoGrande | 83.7% | — |
Math Mistral Medium 3.5 leads
Deepseek Coder v2: 34.9 (#190), Mistral Medium 3.5: 39.1 (#113)
| Benchmark | Deepseek Coder v2 | Mistral Medium 3.5 |
|---|---|---|
| LMArena Math | 1241 | 1431 |
| GSM8K | 94.5% | — |
Knowledge Mistral Medium 3.5 leads
Deepseek Coder v2: 32.3 (#212), Mistral Medium 3.5: 40.0 (#126)
| Benchmark | Deepseek Coder v2 | Mistral Medium 3.5 |
|---|---|---|
| LMArena Expert | 1181 | 1432 |
| ARC (AI2) Challenge | 64.3% | — |
Multimodal Not comparable
Deepseek Coder v2: —, Mistral Medium 3.5: 38.3 (#65)
| Benchmark | Deepseek Coder v2 | Mistral Medium 3.5 |
|---|---|---|
| LMArena Vision | — | 1223 |
Multilingual Mistral Medium 3.5 leads
Deepseek Coder v2: 36.3 (#240), Mistral Medium 3.5: 51.9 (#100)
| Benchmark | Deepseek Coder v2 | Mistral Medium 3.5 |
|---|---|---|
| LMArena Non-English | 1182 | 1404 |
| LMArena Chinese | 1201 | 1442 |
| LMArena French | 1185 | 1448 |
| LMArena German | 1164 | 1451 |
| LMArena Korean | 1104 | 1385 |
| LMArena Russian | 1188 | 1395 |
| LMArena Spanish | 1153 | 1409 |
| LMArena Japanese | 1126 | — |
Instruction Following Mistral Medium 3.5 leads
Deepseek Coder v2: 61.7 (#242), Mistral Medium 3.5: 74.6 (#90)
| Benchmark | Deepseek Coder v2 | Mistral Medium 3.5 |
|---|---|---|
| LMArena Instruction Following | 1180 | 1415 |
Long Context Mistral Medium 3.5 leads
Deepseek Coder v2: 37.0 (#224), Mistral Medium 3.5: 43.2 (#103)
| Benchmark | Deepseek Coder v2 | Mistral Medium 3.5 |
|---|---|---|
| LMArena Longer Query | 1219 | 1415 |
Writing & Preference Mistral Medium 3.5 leads
Deepseek Coder v2: 38.2 (#253), Mistral Medium 3.5: 58.5 (#117)
| Benchmark | Deepseek Coder v2 | Mistral Medium 3.5 |
|---|---|---|
| LMArena Text | 1191 | 1421 |
| LMArena Creative Writing | 1120 | 1374 |
| LMArena Multi-Turn | 1177 | 1423 |
| EQ-Bench 4 | — | 993 |
Frequently asked questions
Is Deepseek Coder v2 better than Mistral Medium 3.5?
Mistral Medium 3.5 is the stronger model overall, scoring 40.2 to 35.9 on the Noometry Index.
Is Deepseek Coder v2 or Mistral Medium 3.5 better for coding?
Deepseek Coder v2 scores higher on coding benchmarks: 38.1 versus 36.0 in the Noometry coding category.
How many benchmarks do Deepseek Coder v2 and Mistral Medium 3.5 share?
16 benchmarks have published results for both models. Deepseek Coder v2 has 24 scored results on Noometry and Mistral Medium 3.5 has 22.