Model comparison

GPT-5.5 Instant vs Qwen3.5 Plus

GPT-5.5 Instant and Qwen3.5 Plus score almost the same on the Noometry Index (42.7 vs 42.9), so choose on price, context window or the category you care about most.

Last verified . 4 shared benchmarks.

GPT-5.5 Instant OpenAI

42.7

Rank #110 Confirmed

Qwen3.5 Plus Alibaba (Qwen)

42.9

Rank #106 Confirmed

Summary

  • They share 4 benchmarks with published results for both. GPT-5.5 Instant scores higher in 2 categories and Qwen3.5 Plus in 2 categories; 3 gaps are clear of the uncertainty.
  • The widest gap is in math, where Qwen3.5 Plus leads 49.6 to 26.5.
  • The biggest single-benchmark swing is OTIS Mock AIME 2024-2025: 68.1% for GPT-5.5 Instant and 86.7% for Qwen3.5 Plus.

Side by side

GPT-5.5 Instant and Qwen3.5 Plus specifications
GPT-5.5 InstantQwen3.5 Plus
ProviderOpenAIAlibaba (Qwen)
Noometry Index42.742.9
Released2026-05-052026-02-16
WeightsProprietaryProprietary
Context window—1M
Max output—66K
Input $ / M tokens—$0.40
Output $ / M tokens—$2.40
Results tracked2715

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding Not comparable

GPT-5.5 Instant: 44.3 (#74), Qwen3.5 Plus: —

Coding benchmarks
BenchmarkGPT-5.5 InstantQwen3.5 Plus
SciCode48.6%—
LMArena Coding1433—
ALE-Bench—621.92

Agentic & Tool Use Not comparable

GPT-5.5 Instant: —, Qwen3.5 Plus: —

Agentic & Tool Use benchmarks
BenchmarkGPT-5.5 InstantQwen3.5 Plus
Vending-Bench 2—0.54

Reasoning Qwen3.5 Plus leads

GPT-5.5 Instant: 24.9 (#155), Qwen3.5 Plus: 32.8 (#74)

Reasoning benchmarks
BenchmarkGPT-5.5 InstantQwen3.5 Plus
Chess Puzzles12%22%
Epoch Capabilities Index142.52146.78
CritPt0%—
LMArena Hard Prompts1426—
Mystery Game Puzzles—17%
DTBench—80.5%
LMCA—36.4%

Math Qwen3.5 Plus leads

GPT-5.5 Instant: 26.5 (#259), Qwen3.5 Plus: 49.6 (#61)

Math benchmarks
BenchmarkGPT-5.5 InstantQwen3.5 Plus
OTIS Mock AIME 2024-202568.1%86.7%
FrontierMath (Tiers 1-3)26.3%—
FrontierMath Tier 42.4%—
LMArena Math1420—
FrontierMath (Feb 2025 set)—21%
FrontierMath Tier 4 (v1)—2.1%

Knowledge GPT-5.5 Instant leads

GPT-5.5 Instant: 48.9 (#74), Qwen3.5 Plus: 46.0 (#83)

Knowledge benchmarks
BenchmarkGPT-5.5 InstantQwen3.5 Plus
GPQA Diamond82.5%84.8%
SimpleQA Verified—25.4%
Vectara Hallucination Rate—10.7%
LMArena Expert1409—

Multimodal Not comparable

GPT-5.5 Instant: 40.0 (#52), Qwen3.5 Plus: —

Multimodal benchmarks
BenchmarkGPT-5.5 InstantQwen3.5 Plus
LMArena Vision1250—
LMArena Document1403—

Multilingual Not comparable

GPT-5.5 Instant: 52.8 (#80), Qwen3.5 Plus: —

Multilingual benchmarks
BenchmarkGPT-5.5 InstantQwen3.5 Plus
LMArena Non-English1417—
LMArena Chinese1456—
LMArena French1428—
LMArena German1411—
LMArena Japanese1408—
LMArena Korean1392—
LMArena Russian1431—
LMArena Spanish1429—

Instruction Following Not comparable

GPT-5.5 Instant: 74.2 (#100), Qwen3.5 Plus: —

Instruction Following benchmarks
BenchmarkGPT-5.5 InstantQwen3.5 Plus
LMArena Instruction Following1406—

Long Context Too close to call

GPT-5.5 Instant: 43.4 (#96), Qwen3.5 Plus: 43.0 (#113)

Long Context benchmarks
BenchmarkGPT-5.5 InstantQwen3.5 Plus
CL-bench—19.8%
CL-bench Life—12.4%
LMArena Longer Query1422—

Writing & Preference Not comparable

GPT-5.5 Instant: 61.8 (#85), Qwen3.5 Plus: —

Writing & Preference benchmarks
BenchmarkGPT-5.5 InstantQwen3.5 Plus
LMArena Text1419—
LMArena Creative Writing1419—
LMArena Multi-Turn1433—

Frequently asked questions

Is GPT-5.5 Instant better than Qwen3.5 Plus?

GPT-5.5 Instant and Qwen3.5 Plus score almost the same on the Noometry Index (42.7 vs 42.9), so choose on price, context window or the category you care about most.

How many benchmarks do GPT-5.5 Instant and Qwen3.5 Plus share?

4 benchmarks have published results for both models. GPT-5.5 Instant has 27 scored results on Noometry and Qwen3.5 Plus has 15.

Related comparisons

Go deeper