Model comparison

GPT-5.4 nano vs Qwen3.5 27B

GPT-5.4 nano and Qwen3.5 27B score almost the same on the Noometry Index (41.9 vs 41.9), so choose on price, context window or the category you care about most.

Last verified . 23 shared benchmarks.

GPT-5.4 nano OpenAI

41.9

Rank #125 Confirmed

Qwen3.5 27B Alibaba (Qwen)

41.9

Rank #127 Confirmed

Summary

  • They share 23 benchmarks with published results for both. GPT-5.4 nano scores higher in 3 categories and Qwen3.5 27B in 6 categories; 9 gaps are clear of the uncertainty.
  • The widest gap is in coding, where GPT-5.4 nano leads 43.6 to 38.9.
  • The biggest single-benchmark swing is WeirdML: 49.2% for GPT-5.4 nano and 39.5% for Qwen3.5 27B.
  • GPT-5.4 nano is cheaper at $0.20 / $1.25 per million input/output tokens, against $0.30 / $2.40 for Qwen3.5 27B.
  • GPT-5.4 nano accepts more context: 400K tokens versus 262K.
  • Qwen3.5 27B has downloadable open weights; the other is API-only.

Side by side

GPT-5.4 nano and Qwen3.5 27B specifications
GPT-5.4 nanoQwen3.5 27B
ProviderOpenAIAlibaba (Qwen)
Noometry Index41.941.9
Released2026-03-172026-02-23
WeightsProprietaryOpen
Context window400K262K
Max output128K66K
Input $ / M tokens$0.20$0.30
Output $ / M tokens$1.25$2.40
Results tracked4028

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding GPT-5.4 nano leads

GPT-5.4 nano: 43.6 (#84), Qwen3.5 27B: 38.9 (#168)

Coding benchmarks
BenchmarkGPT-5.4 nanoQwen3.5 27B
WeirdML49.2%39.5%
LMArena Coding14051427
ALE-Bench1,005349.45
LMArena WebDev—1358
SciCode46.9%—

Agentic & Tool Use Not comparable

GPT-5.4 nano: —, Qwen3.5 27B: —

Agentic & Tool Use benchmarks
BenchmarkGPT-5.4 nanoQwen3.5 27B
Vending-Bench 2—201.98

Reasoning Qwen3.5 27B leads

GPT-5.4 nano: 23.7 (#173), Qwen3.5 27B: 27.5 (#117)

Reasoning benchmarks
BenchmarkGPT-5.4 nanoQwen3.5 27B
LMArena Hard Prompts13811414
DTBench80.3%82.4%
LMCA36.9%34%
ARC-AGI-25.7%—
Kagi LLM Benchmark39.7%—
NYT Connections (extended)—47.9%
ARC-AGI-151.5%—
CritPt9.3%—
Chess Puzzles30%—
Thematic Generalization—45.5%
Mystery Game Puzzles9%—
Epoch Capabilities Index145.81—
ForecastBench57.3—

Math GPT-5.4 nano leads

GPT-5.4 nano: 40.9 (#88), Qwen3.5 27B: 38.8 (#127)

Knowledge GPT-5.4 nano leads

GPT-5.4 nano: 41.9 (#103), Qwen3.5 27B: 38.0 (#150)

Knowledge benchmarks
BenchmarkGPT-5.4 nanoQwen3.5 27B
Vectara Hallucination Rate3.1%12.1%
LMArena Expert13961428
GPQA Diamond78.5%—
SimpleQA Verified11.7%—

Multimodal Qwen3.5 27B leads

GPT-5.4 nano: 36.7 (#78), Qwen3.5 27B: 39.4 (#59)

Multimodal benchmarks
BenchmarkGPT-5.4 nanoQwen3.5 27B
LMArena Vision11961241

Multilingual Qwen3.5 27B leads

GPT-5.4 nano: 48.6 (#140), Qwen3.5 27B: 50.8 (#115)

Multilingual benchmarks
BenchmarkGPT-5.4 nanoQwen3.5 27B
LMArena Non-English13591390
LMArena Chinese13921478
LMArena French13961410
LMArena German13671393
LMArena Japanese13431345
LMArena Korean13201358
LMArena Russian13631390
LMArena Spanish13711407

Instruction Following Qwen3.5 27B leads

GPT-5.4 nano: 71.9 (#144), Qwen3.5 27B: 73.5 (#119)

Instruction Following benchmarks
BenchmarkGPT-5.4 nanoQwen3.5 27B
LMArena Instruction Following13621393

Long Context Qwen3.5 27B leads

GPT-5.4 nano: 41.6 (#137), Qwen3.5 27B: 43.1 (#106)

Long Context benchmarks
BenchmarkGPT-5.4 nanoQwen3.5 27B
LMArena Longer Query13661413

Writing & Preference Qwen3.5 27B leads

GPT-5.4 nano: 55.7 (#142), Qwen3.5 27B: 59.3 (#111)

Writing & Preference benchmarks
BenchmarkGPT-5.4 nanoQwen3.5 27B
LMArena Text13721409
LMArena Creative Writing13141362
LMArena Multi-Turn13821410

Frequently asked questions

Is GPT-5.4 nano better than Qwen3.5 27B?

GPT-5.4 nano and Qwen3.5 27B score almost the same on the Noometry Index (41.9 vs 41.9), so choose on price, context window or the category you care about most.

Which is cheaper, GPT-5.4 nano or Qwen3.5 27B?

GPT-5.4 nano is cheaper. It lists at $0.20 per million input tokens and $1.25 per million output tokens; Qwen3.5 27B lists at $0.30 and $2.40.

Is GPT-5.4 nano or Qwen3.5 27B better for coding?

GPT-5.4 nano scores higher on coding benchmarks: 43.6 versus 38.9 in the Noometry coding category.

Which has the bigger context window?

GPT-5.4 nano does, with 400K tokens against 262K.

How many benchmarks do GPT-5.4 nano and Qwen3.5 27B share?

23 benchmarks have published results for both models. GPT-5.4 nano has 40 scored results on Noometry and Qwen3.5 27B has 28.

Related comparisons

Go deeper