Model comparison

Qwen3.5 397B-A17B vs Qwen3.6 27B

Qwen3.5 397B-A17B is the stronger model overall, scoring 46.0 to 42.2 on the Noometry Index.

Last verified . 8 shared benchmarks.

Qwen3.5 397B-A17B Alibaba (Qwen)

46.0

Rank #67 Confirmed

Qwen3.6 27B Alibaba (Qwen)

42.2

Rank #117 Confirmed

Summary

  • They share 8 benchmarks with published results for both. Qwen3.5 397B-A17B scores higher in 4 categories and Qwen3.6 27B in 1 category; 4 gaps are clear of the uncertainty.
  • The widest gap is in writing & preference, where Qwen3.5 397B-A17B leads 62.3 to 50.3.
  • The biggest single-benchmark swing is Mystery Game Puzzles: 18% for Qwen3.5 397B-A17B and 7% for Qwen3.6 27B.
  • Both cost about the same: $0.60 input and $3.60 output per million tokens.

Side by side

Qwen3.5 397B-A17B and Qwen3.6 27B specifications
Qwen3.5 397B-A17BQwen3.6 27B
ProviderAlibaba (Qwen)Alibaba (Qwen)
Noometry Index46.042.2
Released2026-02-012026-04-22
WeightsOpenOpen
Context window262K262K
Max output66K66K
Input $ / M tokens$0.60$0.60
Output $ / M tokens$3.60$3.60
Results tracked3611

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Qwen3.5 397B-A17B leads

Qwen3.5 397B-A17B: 42.0 (#114), Qwen3.6 27B: 39.1 (#163)

Coding benchmarks
BenchmarkQwen3.5 397B-A17BQwen3.6 27B
LMArena WebDev1400—
SciCode—37.3%
LMArena Coding1465—

Agentic & Tool Use Not comparable

Qwen3.5 397B-A17B: 33.3 (#53), Qwen3.6 27B: —

Agentic & Tool Use benchmarks
BenchmarkQwen3.5 397B-A17BQwen3.6 27B
APEX-Agents24.9%—
τ²-bench Airline81.5%—
τ²-bench Banking9.8%—
τ²-bench Retail84.4%—
τ²-bench Telecom97.8%—

Reasoning Qwen3.5 397B-A17B leads

Qwen3.5 397B-A17B: 34.5 (#70), Qwen3.6 27B: 25.0 (#153)

Reasoning benchmarks
BenchmarkQwen3.5 397B-A17BQwen3.6 27B
Chess Puzzles13%22%
Mystery Game Puzzles18%7%
DTBench87.5%78.1%
LMCA37.9%34.5%
Epoch Capabilities Index146.65146.5
Kagi LLM Benchmark73.7%—
NYT Connections (extended)58.9%—
CritPt—0.9%
Thematic Generalization65.1%—
LMArena Hard Prompts1448—

Math Qwen3.6 27B leads

Qwen3.5 397B-A17B: 46.1 (#73), Qwen3.6 27B: 48.5 (#62)

Math benchmarks
BenchmarkQwen3.5 397B-A17BQwen3.6 27B
FrontierMath (Tiers 1-3)31.2%35.1%
OTIS Mock AIME 2024-202588.9%91.1%
LMArena Math1454—

Knowledge Too close to call

Qwen3.5 397B-A17B: 53.3 (#58), Qwen3.6 27B: 52.4 (#63)

Knowledge benchmarks
BenchmarkQwen3.5 397B-A17BQwen3.6 27B
GPQA Diamond86.4%85.9%
LMArena Expert1462—

Multimodal Not comparable

Qwen3.5 397B-A17B: 40.7 (#44), Qwen3.6 27B: —

Multimodal benchmarks
BenchmarkQwen3.5 397B-A17BQwen3.6 27B
LMArena Vision1263—

Multilingual Not comparable

Qwen3.5 397B-A17B: 53.7 (#59), Qwen3.6 27B: —

Multilingual benchmarks
BenchmarkQwen3.5 397B-A17BQwen3.6 27B
LMArena Non-English1430—
LMArena Chinese1500—
LMArena French1461—
LMArena German1447—
LMArena Japanese1426—
LMArena Korean1384—
LMArena Russian1429—
LMArena Spanish1441—

Instruction Following Not comparable

Qwen3.5 397B-A17B: 75.0 (#77), Qwen3.6 27B: —

Instruction Following benchmarks
BenchmarkQwen3.5 397B-A17BQwen3.6 27B
LMArena Instruction Following1424—

Long Context Not comparable

Qwen3.5 397B-A17B: 44.1 (#74), Qwen3.6 27B: —

Long Context benchmarks
BenchmarkQwen3.5 397B-A17BQwen3.6 27B
LMArena Longer Query1442—

Writing & Preference Qwen3.5 397B-A17B leads

Qwen3.5 397B-A17B: 62.3 (#79), Qwen3.6 27B: 50.3 (#181)

Writing & Preference benchmarks
BenchmarkQwen3.5 397B-A17BQwen3.6 27B
LMArena Text1438—
LMArena Creative Writing1401—
EQ-Bench Creative Writing1478—
EQ-Bench 4—1026
LMArena Multi-Turn1446—

Frequently asked questions

Is Qwen3.5 397B-A17B better than Qwen3.6 27B?

Qwen3.5 397B-A17B is the stronger model overall, scoring 46.0 to 42.2 on the Noometry Index.

Which is cheaper, Qwen3.5 397B-A17B or Qwen3.6 27B?

Qwen3.6 27B is cheaper. It lists at $0.60 per million input tokens and $3.60 per million output tokens; Qwen3.5 397B-A17B lists at $0.60 and $3.60.

Is Qwen3.5 397B-A17B or Qwen3.6 27B better for coding?

Qwen3.5 397B-A17B scores higher on coding benchmarks: 42.0 versus 39.1 in the Noometry coding category.

Which has the bigger context window?

Both accept 262K tokens.

How many benchmarks do Qwen3.5 397B-A17B and Qwen3.6 27B share?

8 benchmarks have published results for both models. Qwen3.5 397B-A17B has 36 scored results on Noometry and Qwen3.6 27B has 11.

Related comparisons

Go deeper