Model comparison

GPT-5.3 Chat vs Qwen3.5 Plus

GPT-5.3 Chat and Qwen3.5 Plus score almost the same on the Noometry Index (42.8 vs 42.9), so choose on price, context window or the category you care about most.

Last verified . 0 shared benchmarks.

GPT-5.3 Chat OpenAI

42.8

Rank #109 Confirmed

Qwen3.5 Plus Alibaba (Qwen)

42.9

Rank #106 Confirmed

Summary

  • The widest gap is in math, where Qwen3.5 Plus leads 49.6 to 38.2.
  • Qwen3.5 Plus is cheaper at $0.40 / $2.40 per million input/output tokens, against $1.75 / $14 for GPT-5.3 Chat.
  • Qwen3.5 Plus accepts more context: 1M tokens versus 128K.

Side by side

GPT-5.3 Chat and Qwen3.5 Plus specifications
GPT-5.3 ChatQwen3.5 Plus
ProviderOpenAIAlibaba (Qwen)
Noometry Index42.842.9
Released2026-03-032026-02-16
WeightsProprietaryProprietary
Context window128K1M
Max output16K66K
Input $ / M tokens$1.75$0.40
Output $ / M tokens$14$2.40
Results tracked1815

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Not comparable

GPT-5.3 Chat: 41.4 (#124), Qwen3.5 Plus: —

Coding benchmarks
BenchmarkGPT-5.3 ChatQwen3.5 Plus
LMArena Coding1408—
ALE-Bench—621.92

Agentic & Tool Use Not comparable

GPT-5.3 Chat: —, Qwen3.5 Plus: —

Agentic & Tool Use benchmarks
BenchmarkGPT-5.3 ChatQwen3.5 Plus
Vending-Bench 2—0.54

Reasoning Qwen3.5 Plus leads

GPT-5.3 Chat: 28.5 (#102), Qwen3.5 Plus: 32.8 (#74)

Reasoning benchmarks
BenchmarkGPT-5.3 ChatQwen3.5 Plus
Chess Puzzles—22%
LMArena Hard Prompts1399—
Mystery Game Puzzles—17%
DTBench—80.5%
LMCA—36.4%
Epoch Capabilities Index—146.78

Math Qwen3.5 Plus leads

GPT-5.3 Chat: 38.2 (#142), Qwen3.5 Plus: 49.6 (#61)

Math benchmarks
BenchmarkGPT-5.3 ChatQwen3.5 Plus
OTIS Mock AIME 2024-2025—86.7%
LMArena Math1389—
FrontierMath (Feb 2025 set)—21%
FrontierMath Tier 4 (v1)—2.1%

Knowledge Qwen3.5 Plus leads

GPT-5.3 Chat: 38.8 (#140), Qwen3.5 Plus: 46.0 (#83)

Knowledge benchmarks
BenchmarkGPT-5.3 ChatQwen3.5 Plus
GPQA Diamond—84.8%
SimpleQA Verified—25.4%
Vectara Hallucination Rate—10.7%
LMArena Expert1397—

Multilingual Not comparable

GPT-5.3 Chat: 50.3 (#124), Qwen3.5 Plus: —

Multilingual benchmarks
BenchmarkGPT-5.3 ChatQwen3.5 Plus
LMArena Non-English1382—
LMArena Chinese1432—
LMArena French1397—
LMArena German1384—
LMArena Japanese1352—
LMArena Korean1346—
LMArena Russian1400—
LMArena Spanish1371—

Instruction Following Not comparable

GPT-5.3 Chat: 72.8 (#129), Qwen3.5 Plus: —

Instruction Following benchmarks
BenchmarkGPT-5.3 ChatQwen3.5 Plus
LMArena Instruction Following1378—

Long Context Too close to call

GPT-5.3 Chat: 42.6 (#120), Qwen3.5 Plus: 43.0 (#113)

Long Context benchmarks
BenchmarkGPT-5.3 ChatQwen3.5 Plus
CL-bench—19.8%
CL-bench Life—12.4%
LMArena Longer Query1396—

Writing & Preference Not comparable

GPT-5.3 Chat: 63.1 (#68), Qwen3.5 Plus: —

Writing & Preference benchmarks
BenchmarkGPT-5.3 ChatQwen3.5 Plus
LMArena Text1389—
LMArena Creative Writing1355—
EQ-Bench Creative Writing1690—
LMArena Multi-Turn1412—

Frequently asked questions

Is GPT-5.3 Chat better than Qwen3.5 Plus?

GPT-5.3 Chat and Qwen3.5 Plus score almost the same on the Noometry Index (42.8 vs 42.9), so choose on price, context window or the category you care about most.

Which is cheaper, GPT-5.3 Chat or Qwen3.5 Plus?

Qwen3.5 Plus is cheaper. It lists at $0.40 per million input tokens and $2.40 per million output tokens; GPT-5.3 Chat lists at $1.75 and $14.

Which has the bigger context window?

Qwen3.5 Plus does, with 1M tokens against 128K.

How many benchmarks do GPT-5.3 Chat and Qwen3.5 Plus share?

0 benchmarks have published results for both models. GPT-5.3 Chat has 18 scored results on Noometry and Qwen3.5 Plus has 15.

Related comparisons

Go deeper