Model comparison

GPT-5.3 Codex vs Qwen3.7 Plus

GPT-5.3 Codex and Qwen3.7 Plus score almost the same on the Noometry Index (45.8 vs 45.3), so choose on price, context window or the category you care about most.

Last verified . 1 shared benchmarks.

GPT-5.3 Codex OpenAI

45.8

Rank #69 Reported

Qwen3.7 Plus Alibaba (Qwen)

45.3

Rank #72 Confirmed

Summary

  • They share 1 benchmark with published results for both. GPT-5.3 Codex scores higher in 2 categories and Qwen3.7 Plus in 0 categories; 2 gaps are clear of the uncertainty.
  • The widest gap is in agentic & tool use, where GPT-5.3 Codex leads 48.0 to 21.4.
  • Qwen3.7 Plus is cheaper at $0.40 / $1.60 per million input/output tokens, against $1.75 / $14 for GPT-5.3 Codex.
  • Qwen3.7 Plus accepts more context: 1M tokens versus 400K.

Side by side

GPT-5.3 Codex and Qwen3.7 Plus specifications
GPT-5.3 CodexQwen3.7 Plus
ProviderOpenAIAlibaba (Qwen)
Noometry Index45.845.3
Released2026-02-052026-06-02
WeightsProprietaryProprietary
Context window400K1M
Max output128K131K
Input $ / M tokens$1.75$0.40
Output $ / M tokens$14$1.60
Results tracked832

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding GPT-5.3 Codex leads

GPT-5.3 Codex: 48.6 (#56), Qwen3.7 Plus: 36.6 (#206)

Coding benchmarks
BenchmarkGPT-5.3 CodexQwen3.7 Plus
SWE-bench Verified74.8%—
FrontierCode—10.2%
LMArena WebDev1409—
SciCode—45.5%
WeirdML79.3%—
LMArena Coding—1473
ALE-Bench1,655—

Agentic & Tool Use GPT-5.3 Codex leads

GPT-5.3 Codex: 48.0 (#9), Qwen3.7 Plus: 21.4 (#138)

Agentic & Tool Use benchmarks
BenchmarkGPT-5.3 CodexQwen3.7 Plus
Terminal-Bench78.4%—
OSWorld 2.0—2.8%
METR Time Horizons74.5%—
Vending-Bench 25,940—

Reasoning Not comparable

GPT-5.3 Codex: —, Qwen3.7 Plus: 39.3 (#59)

Reasoning benchmarks
BenchmarkGPT-5.3 CodexQwen3.7 Plus
Epoch Capabilities Index156.77147.37
NYT Connections (extended)—74.8%
CritPt—9.1%
Chess Puzzles—24%
LMArena Hard Prompts—1460
Mystery Game Puzzles—17%
DTBench—84%
LMCA—37.6%

Math Not comparable

GPT-5.3 Codex: —, Qwen3.7 Plus: 50.5 (#56)

Math benchmarks
BenchmarkGPT-5.3 CodexQwen3.7 Plus
FrontierMath (Tiers 1-3)—34.4%
OTIS Mock AIME 2024-2025—93.3%
LMArena Math—1466

Knowledge Not comparable

GPT-5.3 Codex: —, Qwen3.7 Plus: 54.9 (#51)

Knowledge benchmarks
BenchmarkGPT-5.3 CodexQwen3.7 Plus
GPQA Diamond—87.9%
LMArena Expert—1467

Multimodal Not comparable

GPT-5.3 Codex: —, Qwen3.7 Plus: 41.8 (#33)

Multimodal benchmarks
BenchmarkGPT-5.3 CodexQwen3.7 Plus
LMArena Vision—1279
LMArena Document—1444

Multilingual Not comparable

GPT-5.3 Codex: —, Qwen3.7 Plus: 54.8 (#38)

Multilingual benchmarks
BenchmarkGPT-5.3 CodexQwen3.7 Plus
LMArena Non-English—1445
LMArena Chinese—1510
LMArena French—1473
LMArena German—1471
LMArena Japanese—1413
LMArena Korean—1415
LMArena Russian—1457
LMArena Spanish—1457

Instruction Following Not comparable

GPT-5.3 Codex: —, Qwen3.7 Plus: 75.8 (#52)

Instruction Following benchmarks
BenchmarkGPT-5.3 CodexQwen3.7 Plus
LMArena Instruction Following—1440

Long Context Not comparable

GPT-5.3 Codex: —, Qwen3.7 Plus: 44.5 (#65)

Long Context benchmarks
BenchmarkGPT-5.3 CodexQwen3.7 Plus
LMArena Longer Query—1455

Writing & Preference Not comparable

GPT-5.3 Codex: —, Qwen3.7 Plus: 64.3 (#56)

Writing & Preference benchmarks
BenchmarkGPT-5.3 CodexQwen3.7 Plus
LMArena Text—1455
LMArena Creative Writing—1439
LMArena Multi-Turn—1460

Frequently asked questions

Is GPT-5.3 Codex better than Qwen3.7 Plus?

GPT-5.3 Codex and Qwen3.7 Plus score almost the same on the Noometry Index (45.8 vs 45.3), so choose on price, context window or the category you care about most.

Which is cheaper, GPT-5.3 Codex or Qwen3.7 Plus?

Qwen3.7 Plus is cheaper. It lists at $0.40 per million input tokens and $1.60 per million output tokens; GPT-5.3 Codex lists at $1.75 and $14.

Is GPT-5.3 Codex or Qwen3.7 Plus better for coding?

GPT-5.3 Codex scores higher on coding benchmarks: 48.6 versus 36.6 in the Noometry coding category.

Which has the bigger context window?

Qwen3.7 Plus does, with 1M tokens against 400K.

How many benchmarks do GPT-5.3 Codex and Qwen3.7 Plus share?

1 benchmark has published results for both models. GPT-5.3 Codex has 8 scored results on Noometry and Qwen3.7 Plus has 32.

Related comparisons

Go deeper