Model comparison

DeepSeek-V3.1 vs Qwen3.5 Plus

DeepSeek-V3.1 and Qwen3.5 Plus score almost the same on the Noometry Index (42.8 vs 42.9), so choose on price, context window or the category you care about most.

Last verified . 4 shared benchmarks.

DeepSeek-V3.1 DeepSeek

42.8

Rank #108 Confirmed

Qwen3.5 Plus Alibaba (Qwen)

42.9

Rank #106 Confirmed

Summary

  • They share 4 benchmarks with published results for both. DeepSeek-V3.1 scores higher in 0 categories and Qwen3.5 Plus in 4 categories; 4 gaps are clear of the uncertainty.
  • The widest gap is in math, where Qwen3.5 Plus leads 49.6 to 38.9.
  • The biggest single-benchmark swing is LMCA: 24.3% for DeepSeek-V3.1 and 36.4% for Qwen3.5 Plus.
  • DeepSeek-V3.1 is cheaper at $0.25 / $0.95 per million input/output tokens, against $0.40 / $2.40 for Qwen3.5 Plus.
  • Qwen3.5 Plus accepts more context: 1M tokens versus 164K.
  • DeepSeek-V3.1 has downloadable open weights; the other is API-only.

Side by side

DeepSeek-V3.1 and Qwen3.5 Plus specifications
DeepSeek-V3.1Qwen3.5 Plus
ProviderDeepSeekAlibaba (Qwen)
Noometry Index42.842.9
Released2025-08-212026-02-16
WeightsOpenProprietary
Context window164K1M
Max output8K66K
Input $ / M tokens$0.25$0.40
Output $ / M tokens$0.95$2.40
Results tracked2715

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Not comparable

DeepSeek-V3.1: 40.3 (#144), Qwen3.5 Plus: —

Coding benchmarks
BenchmarkDeepSeek-V3.1Qwen3.5 Plus
WeirdML38.4%—
LMArena Coding1417—
ALE-Bench—621.92

Agentic & Tool Use Not comparable

DeepSeek-V3.1: —, Qwen3.5 Plus: —

Agentic & Tool Use benchmarks
BenchmarkDeepSeek-V3.1Qwen3.5 Plus
Vending-Bench 2—0.54

Reasoning Qwen3.5 Plus leads

DeepSeek-V3.1: 27.9 (#110), Qwen3.5 Plus: 32.8 (#74)

Reasoning benchmarks
BenchmarkDeepSeek-V3.1Qwen3.5 Plus
DTBench82.7%80.5%
LMCA24.3%36.4%
Epoch Capabilities Index139.92146.78
SimpleBench40%—
Kagi LLM Benchmark53.2%—
Chess Puzzles—22%
LMArena Hard Prompts1417—
Mystery Game Puzzles—17%
ForecastBench58—

Math Qwen3.5 Plus leads

DeepSeek-V3.1: 38.9 (#122), Qwen3.5 Plus: 49.6 (#61)

Math benchmarks
BenchmarkDeepSeek-V3.1Qwen3.5 Plus
OTIS Mock AIME 2024-2025—86.7%
LMArena Math1420—
FrontierMath (Feb 2025 set)—21%
FrontierMath Tier 4 (v1)—2.1%

Knowledge Qwen3.5 Plus leads

DeepSeek-V3.1: 43.7 (#90), Qwen3.5 Plus: 46.0 (#83)

Knowledge benchmarks
BenchmarkDeepSeek-V3.1Qwen3.5 Plus
Vectara Hallucination Rate5.5%10.7%
GPQA Diamond—84.8%
SimpleQA Verified—25.4%
LMArena Expert1405—

Multilingual Not comparable

DeepSeek-V3.1: 51.6 (#106), Qwen3.5 Plus: —

Multilingual benchmarks
BenchmarkDeepSeek-V3.1Qwen3.5 Plus
LMArena Non-English1400—
LMArena Chinese1469—
LMArena French1447—
LMArena German1411—
LMArena Japanese1378—
LMArena Korean1337—
LMArena Russian1405—
LMArena Spanish1431—

Instruction Following Not comparable

DeepSeek-V3.1: 73.9 (#110), Qwen3.5 Plus: —

Instruction Following benchmarks
BenchmarkDeepSeek-V3.1Qwen3.5 Plus
LMArena Instruction Following1400—

Long Context Qwen3.5 Plus leads

DeepSeek-V3.1: 36.3 (#232), Qwen3.5 Plus: 43.0 (#113)

Long Context benchmarks
BenchmarkDeepSeek-V3.1Qwen3.5 Plus
Fiction.LiveBench52.8%—
CL-bench—19.8%
CL-bench Life—12.4%
LMArena Longer Query1422—

Writing & Preference Not comparable

DeepSeek-V3.1: 60.3 (#98), Qwen3.5 Plus: —

Writing & Preference benchmarks
BenchmarkDeepSeek-V3.1Qwen3.5 Plus
LMArena Text1420—
LMArena Creative Writing1401—
EQ-Bench Creative Writing1436—
LMArena Multi-Turn1408—

Frequently asked questions

Is DeepSeek-V3.1 better than Qwen3.5 Plus?

DeepSeek-V3.1 and Qwen3.5 Plus score almost the same on the Noometry Index (42.8 vs 42.9), so choose on price, context window or the category you care about most.

Which is cheaper, DeepSeek-V3.1 or Qwen3.5 Plus?

DeepSeek-V3.1 is cheaper. It lists at $0.25 per million input tokens and $0.95 per million output tokens; Qwen3.5 Plus lists at $0.40 and $2.40.

Which has the bigger context window?

Qwen3.5 Plus does, with 1M tokens against 164K.

How many benchmarks do DeepSeek-V3.1 and Qwen3.5 Plus share?

4 benchmarks have published results for both models. DeepSeek-V3.1 has 27 scored results on Noometry and Qwen3.5 Plus has 15.

Related comparisons

Go deeper