Model comparison

GPT-5 Mini vs Qwen3.5 27B

GPT-5 Mini and Qwen3.5 27B score almost the same on the Noometry Index (41.8 vs 41.9), so choose on price, context window or the category you care about most.

Last verified . 24 shared benchmarks.

GPT-5 Mini OpenAI

41.8

Rank #128 Confirmed

Qwen3.5 27B Alibaba (Qwen)

41.9

Rank #127 Confirmed

Summary

  • They share 24 benchmarks with published results for both. GPT-5 Mini scores higher in 4 categories and Qwen3.5 27B in 5 categories; 9 gaps are clear of the uncertainty.
  • The widest gap is in math, where GPT-5 Mini leads 46.7 to 38.8.
  • The biggest single-benchmark swing is WeirdML: 52.7% for GPT-5 Mini and 39.5% for Qwen3.5 27B.
  • GPT-5 Mini is cheaper at $0.25 / $2 per million input/output tokens, against $0.30 / $2.40 for Qwen3.5 27B.
  • GPT-5 Mini accepts more context: 400K tokens versus 262K.
  • Qwen3.5 27B has downloadable open weights; the other is API-only.

Side by side

GPT-5 Mini and Qwen3.5 27B specifications
GPT-5 MiniQwen3.5 27B
ProviderOpenAIAlibaba (Qwen)
Noometry Index41.841.9
Released2025-08-072026-02-23
WeightsProprietaryOpen
Context window400K262K
Max output128K66K
Input $ / M tokens$0.25$0.30
Output $ / M tokens$2$2.40
Results tracked6028

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding GPT-5 Mini leads

GPT-5 Mini: 40.1 (#146), Qwen3.5 27B: 38.9 (#168)

Coding benchmarks
BenchmarkGPT-5 MiniQwen3.5 27B
WeirdML52.7%39.5%
LMArena Coding14061427
ALE-Bench799.77349.45
SWE-bench Verified64.7%—
SWE-bench Verified (bash only)59.8%—
LMArena WebDev—1358
SWE-bench Multilingual39.7%—
SciCode39.2%—
AlgoTune1.38—

Agentic & Tool Use Not comparable

GPT-5 Mini: 31.1 (#70), Qwen3.5 27B: —

Agentic & Tool Use benchmarks
BenchmarkGPT-5 MiniQwen3.5 27B
Vending-Bench 2-31.18201.98
Terminal-Bench34.8%—
Berkeley Function Calling Leaderboard55.5%—

Reasoning Qwen3.5 27B leads

GPT-5 Mini: 23.9 (#168), Qwen3.5 27B: 27.5 (#117)

Reasoning benchmarks
BenchmarkGPT-5 MiniQwen3.5 27B
LMArena Hard Prompts13801414
DTBench80.5%82.4%
LMCA34.2%34%
ARC-AGI-24.4%—
Kagi LLM Benchmark70.3%—
NYT Connections (extended)—47.9%
ARC-AGI-154.3%—
CritPt0%—
Chess Puzzles30%—
EnigmaEval8.2%—
Thematic Generalization—45.5%
Mystery Game Puzzles10%—
Epoch Capabilities Index145.52—
ForecastBench61—

Math GPT-5 Mini leads

GPT-5 Mini: 46.7 (#69), Qwen3.5 27B: 38.8 (#127)

Knowledge GPT-5 Mini leads

GPT-5 Mini: 45.6 (#86), Qwen3.5 27B: 38.0 (#150)

Knowledge benchmarks
BenchmarkGPT-5 MiniQwen3.5 27B
Vectara Hallucination Rate12.9%12.1%
LMArena Expert13791428
GPQA Diamond75%—
Humanity's Last Exam19.4%—
SimpleQA Verified21.6%—
MMLU-Pro83.5%—
Confabulations13.3%—
GPQA (HELM)75.6%—

Multimodal Qwen3.5 27B leads

GPT-5 Mini: 35.6 (#85), Qwen3.5 27B: 39.4 (#59)

Multimodal benchmarks
BenchmarkGPT-5 MiniQwen3.5 27B
LMArena Vision12021241
VPCT40.2%—

Multilingual Qwen3.5 27B leads

GPT-5 Mini: 48.9 (#137), Qwen3.5 27B: 50.8 (#115)

Multilingual benchmarks
BenchmarkGPT-5 MiniQwen3.5 27B
LMArena Non-English13631390
LMArena Chinese13851478
LMArena French13861410
LMArena German13661393
LMArena Japanese13411345
LMArena Korean13081358
LMArena Russian13621390
LMArena Spanish13551407

Instruction Following GPT-5 Mini leads

GPT-5 Mini: 76.2 (#46), Qwen3.5 27B: 73.5 (#119)

Instruction Following benchmarks
BenchmarkGPT-5 MiniQwen3.5 27B
LMArena Instruction Following13571393
IFEval92.7%—

Long Context Qwen3.5 27B leads

GPT-5 Mini: 41.9 (#132), Qwen3.5 27B: 43.1 (#106)

Long Context benchmarks
BenchmarkGPT-5 MiniQwen3.5 27B
LMArena Longer Query13551413
Fiction.LiveBench69.4%—

Writing & Preference Qwen3.5 27B leads

GPT-5 Mini: 55.2 (#148), Qwen3.5 27B: 59.3 (#111)

Writing & Preference benchmarks
BenchmarkGPT-5 MiniQwen3.5 27B
LMArena Text13731409
LMArena Creative Writing13251362
LMArena Multi-Turn13631410
Short-Story Creative Writing83.1%—
EQ-Bench Creative Writing1313—
WildBench85.5%—

Frequently asked questions

Is GPT-5 Mini better than Qwen3.5 27B?

GPT-5 Mini and Qwen3.5 27B score almost the same on the Noometry Index (41.8 vs 41.9), so choose on price, context window or the category you care about most.

Which is cheaper, GPT-5 Mini or Qwen3.5 27B?

GPT-5 Mini is cheaper. It lists at $0.25 per million input tokens and $2 per million output tokens; Qwen3.5 27B lists at $0.30 and $2.40.

Is GPT-5 Mini or Qwen3.5 27B better for coding?

GPT-5 Mini scores higher on coding benchmarks: 40.1 versus 38.9 in the Noometry coding category.

Which has the bigger context window?

GPT-5 Mini does, with 400K tokens against 262K.

How many benchmarks do GPT-5 Mini and Qwen3.5 27B share?

24 benchmarks have published results for both models. GPT-5 Mini has 60 scored results on Noometry and Qwen3.5 27B has 28.

Related comparisons

Go deeper