Model comparison

GPT-5.2 Codex vs Step 3.5 Flash

GPT-5.2 Codex and Step 3.5 Flash score almost the same on the Noometry Index (42.6 vs 42.3), so choose on price, context window or the category you care about most.

Last verified . 0 shared benchmarks.

GPT-5.2 Codex OpenAI

42.6

Rank #111 Reported

Step 3.5 Flash StepFun

42.3

Rank #116 Confirmed

Summary

  • The widest gap is in coding, where GPT-5.2 Codex leads 45.5 to 42.4.
  • Step 3.5 Flash is cheaper at $0.10 / $0.30 per million input/output tokens, against $1.75 / $14 for GPT-5.2 Codex.
  • GPT-5.2 Codex accepts more context: 400K tokens versus 256K.
  • Step 3.5 Flash has downloadable open weights; the other is API-only.

Side by side

GPT-5.2 Codex and Step 3.5 Flash specifications
GPT-5.2 CodexStep 3.5 Flash
ProviderOpenAIStepFun
Noometry Index42.642.3
Released2025-12-182026-01-29
WeightsProprietaryOpen
Context window400K256K
Max output128K256K
Input $ / M tokens$1.75$0.10
Output $ / M tokens$14$0.30
Results tracked519

Sponsored placements are available on pages like this one. Advertise on Noometry

Category by category

Coding GPT-5.2 Codex leads

GPT-5.2 Codex: 45.5 (#71), Step 3.5 Flash: 42.4 (#105)

Coding benchmarks
BenchmarkGPT-5.2 CodexStep 3.5 Flash
SWE-bench Verified (bash only)72.8%—
LMArena WebDev1339—
SWE-bench Multilingual66.3%—
LMArena Coding—1436
ALE-Bench1,300—

Agentic & Tool Use Not comparable

GPT-5.2 Codex: 41.0 (#22), Step 3.5 Flash: —

Agentic & Tool Use benchmarks
BenchmarkGPT-5.2 CodexStep 3.5 Flash
Terminal-Bench66.5%—

Reasoning Not comparable

GPT-5.2 Codex: —, Step 3.5 Flash: 22.2 (#202)

Reasoning benchmarks
BenchmarkGPT-5.2 CodexStep 3.5 Flash
NYT Connections (extended)—28.4%
LMArena Hard Prompts—1411

Math Not comparable

GPT-5.2 Codex: —, Step 3.5 Flash: 42.6 (#84)

Math benchmarks
BenchmarkGPT-5.2 CodexStep 3.5 Flash
MathArena Final-Answer Competitions—66.8%
LMArena Math—1408

Knowledge Not comparable

GPT-5.2 Codex: —, Step 3.5 Flash: 39.6 (#132)

Knowledge benchmarks
BenchmarkGPT-5.2 CodexStep 3.5 Flash
LMArena Expert—1421

Multilingual Not comparable

GPT-5.2 Codex: —, Step 3.5 Flash: 50.5 (#119)

Multilingual benchmarks
BenchmarkGPT-5.2 CodexStep 3.5 Flash
LMArena Non-English—1385
LMArena Chinese—1447
LMArena French—1421
LMArena German—1405
LMArena Japanese—1354
LMArena Korean—1352
LMArena Russian—1385
LMArena Spanish—1419

Instruction Following Not comparable

GPT-5.2 Codex: —, Step 3.5 Flash: 73.1 (#124)

Instruction Following benchmarks
BenchmarkGPT-5.2 CodexStep 3.5 Flash
LMArena Instruction Following—1385

Long Context Not comparable

GPT-5.2 Codex: —, Step 3.5 Flash: 42.8 (#117)

Long Context benchmarks
BenchmarkGPT-5.2 CodexStep 3.5 Flash
LMArena Longer Query—1402

Writing & Preference Not comparable

GPT-5.2 Codex: —, Step 3.5 Flash: 58.8 (#113)

Writing & Preference benchmarks
BenchmarkGPT-5.2 CodexStep 3.5 Flash
LMArena Text—1403
LMArena Creative Writing—1357
LMArena Multi-Turn—1405

Frequently asked questions

Is GPT-5.2 Codex better than Step 3.5 Flash?

GPT-5.2 Codex and Step 3.5 Flash score almost the same on the Noometry Index (42.6 vs 42.3), so choose on price, context window or the category you care about most.

Which is cheaper, GPT-5.2 Codex or Step 3.5 Flash?

Step 3.5 Flash is cheaper. It lists at $0.10 per million input tokens and $0.30 per million output tokens; GPT-5.2 Codex lists at $1.75 and $14.

Is GPT-5.2 Codex or Step 3.5 Flash better for coding?

GPT-5.2 Codex scores higher on coding benchmarks: 45.5 versus 42.4 in the Noometry coding category.

Which has the bigger context window?

GPT-5.2 Codex does, with 400K tokens against 256K.

How many benchmarks do GPT-5.2 Codex and Step 3.5 Flash share?

0 benchmarks have published results for both models. GPT-5.2 Codex has 5 scored results on Noometry and Step 3.5 Flash has 19.

Related comparisons

Go deeper