OpenAI, proprietary

GPT-3.5-turbo

GPT-3.5-turbo by OpenAI ranks 350th of 354 ranked models on the Noometry Index as of October 2026, with a score of 23.2. Its strongest category is long context, where it ranks 254th. API pricing starts at $0.50 per million input tokens and $1.50 per million output tokens, with a 16K-token context window.

Last verified

Specifications

Noometry rank
#350 of 354
Index score
23.2
Evidence
Confirmed 44 results
Provider
OpenAI
Released
March 1, 2023
Weights
Proprietary
Reasoning
No
Context window
16K
Max output
4K
Input price
$0.50 / M
Output price
$1.50 / M
Blended price
$0.75 / M
Output speed
Not measured
Value
#133 of 219
Knowledge cutoff
September 2021
Input
text

Category scores

Each category score combines every public result we have in that category.

GPT-3.5-turbo category scores
  1. Coding 23.9
  2. Reasoning 13.8
  3. Math 6.3
  4. Knowledge 10.0
  5. Multilingual 31.5
  6. Instruction Following 57.9
  7. Long Context 34.0
  8. Writing & Preference 25.3
GPT-3.5-turbo category ranks
CategoryScoreRankResults
Coding23.9#3314
Reasoning13.8#3325
Math6.3#3274
Knowledge10.0#3032
Multilingual31.5#2581
Instruction Following57.9#2621
Long Context34.0#2541
Writing & Preference25.3#3054

Strengths and weaknesses

Categories where GPT-3.5-turbo places highest and lowest among the models ranked in each, with its score against that category's median.

Strongest categories

GPT-3.5-turbo: strongest categories
CategoryScorevs medianRank
Long Context34.0−7.0#254 of 296, top 86%
Instruction Following57.9−13.4#262 of 305, top 86%
Multilingual31.5−15.9#258 of 297, top 87%

Weakest categories

GPT-3.5-turbo: weakest categories
CategoryScorevs medianRank
Math6.3−30.3#327 of 327, top 100%
Writing & Preference25.3−28.5#305 of 312, top 98%
Coding23.9−14.8#331 of 340, top 98%

Closest competitors

The models ranked just above and below GPT-3.5-turbo. When scores are this close, price and speed are often the better way to choose.

Models ranked closest to GPT-3.5-turbo
ModelRankScoreBlended $/MSpeed
Claude 2#34625.0——Compare
DeepSeek LLM 67B#34724.9——Compare
Llama 13b#34824.4——Compare
Llama 2-70B#34924.4——Compare
Mistral 7B#35123.0$0.25—Compare
Llama 3.1-8B#35223.0$0.0575—Compare
Gemma 3 1B#35321.1——Compare
Llama 3.2 1B#35420.1$0.0705—Compare

Sponsored placements are available on pages like this one. Advertise on Noometry

Benchmark results

Every published result we track, with its source. Bold rows are the ones used for ranking; where several exist we prefer independent runs over self-reported numbers.

Coding

GPT-3.5-turbo Coding benchmark results
BenchmarkScorePositionSettingSourceDate
WeirdML3.5%#117 of 119, top 99%Epoch AI
BigCodeBench Instruct39.1%#35 of 64, top 55%BigCodeBench2024-01-25
LMArena Coding1136#260 of 294, top 89%LMArena2026-10-08
LMArena Coding1116LMArena2026-10-08
BigCodeBench Complete50.6%#31 of 66, top 47%BigCodeBench2024-01-25
HumanEval+66.5%may 2023EvalPlus
HumanEval+70.7%#20 of 45, top 45%nov 2023EvalPlus
MBPP+69.7%#14 of 38, top 37%nov 2023EvalPlus

Agentic & Tool Use

GPT-3.5-turbo Agentic & Tool Use benchmark results
BenchmarkScorePositionSettingSourceDate
METR Time Horizons21.5%#32 of 32, top 100%Epoch AI

Reasoning

GPT-3.5-turbo Reasoning benchmark results
BenchmarkScorePositionSettingSourceDate
Chess Puzzles0%#115 of 129, top 90%Epoch AI2026-08-07
LMArena Hard Prompts1099LMArena2026-10-08
LMArena Hard Prompts1108#266 of 297, top 90%LMArena2026-10-08
Mystery Game Puzzles3%#73 of 74, top 99%Epoch AI2026-08-27
DTBench48.5%#141 of 151, top 94%Epoch AI
LMCA9.7%#113 of 125, top 91%Epoch AI
Adversarial NLI58.1%Best of 9Epoch AI
BIG-Bench Hard61.6%#12 of 27, top 45%Epoch AI
CommonsenseQA 2.057%Best of 2Epoch AI
Epoch Capabilities Index115.7Epoch AI2024-01-25
Epoch Capabilities Index113.36Epoch AI2023-06-13
Epoch Capabilities Index118.55#174 of 213, top 82%Epoch AI2023-11-06
ForecastBench50.4#72 of 72, top 100%Epoch AI
WinoGrande81.6%#11 of 43, top 26%Epoch AI
WinoGrande68.8%Epoch AI

Math

GPT-3.5-turbo Math benchmark results
BenchmarkScorePositionSettingSourceDate
FrontierMath (Tiers 1-3)0%#81 of 81, top 100%Epoch AI2026-08-27
OTIS Mock AIME 2024-20252.2%#160 of 173, top 93%Epoch AI2026-08-07
LMArena Math1141LMArena2026-10-08
LMArena Math1142#256 of 285, top 90%LMArena2026-10-08
MATH Level 511.6%Epoch AI2025-01-27
MATH Level 515.9%#66 of 79, top 84%Epoch AI2025-01-27
GSM8K57.8%#17 of 38, top 45%Epoch AI

Knowledge

GPT-3.5-turbo Knowledge benchmark results
BenchmarkScorePositionSettingSourceDate
GPQA Diamond28%#173 of 186, top 94%Epoch AI2025-01-27
GPQA Diamond27.2%Epoch AI2025-01-27
LMArena Expert1070#256 of 273, top 94%LMArena2026-10-08
LMArena Expert1066LMArena2026-10-08
ARC (AI2) Challenge87.4%#7 of 39, top 18%Epoch AI
BoolQ87%#5 of 23, top 22%Epoch AI
MMLU71.4%#41 of 81, top 51%Epoch AI
MMLU68.9%Epoch AI
MMLU67.3%Epoch AI
OpenBookQA86%#4 of 19, top 22%Epoch AI
TriviaQA85.8%#4 of 25, top 16%Epoch AI

Multilingual

GPT-3.5-turbo Multilingual benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Non-English1074LMArena2026-10-08
LMArena Non-English1108#258 of 297, top 87%LMArena2026-10-08
LMArena Chinese1075#260 of 285, top 92%LMArena2026-10-08
LMArena Chinese1012LMArena2026-10-08
LMArena French1066LMArena2026-10-08
LMArena French1118#210 of 223, top 95%LMArena2026-10-08
LMArena German1090#212 of 231, top 92%LMArena2026-10-08
LMArena German1056LMArena2026-10-08
LMArena Japanese1043#191 of 211, top 91%LMArena2026-10-08
LMArena Korean1019#197 of 213, top 93%LMArena2026-10-08
LMArena Russian1077LMArena2026-10-08
LMArena Russian1123#252 of 283, top 90%LMArena2026-10-08
LMArena Spanish1093LMArena2026-10-08
LMArena Spanish1121#211 of 226, top 94%LMArena2026-10-08

Instruction Following

GPT-3.5-turbo Instruction Following benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Instruction Following1119#259 of 298, top 87%LMArena2026-10-08
LMArena Instruction Following1093LMArena2026-10-08

Long Context

GPT-3.5-turbo Long Context benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Longer Query1070LMArena2026-10-08
LMArena Longer Query1121#262 of 291, top 91%LMArena2026-10-08

Writing & Preference

GPT-3.5-turbo Writing & Preference benchmark results
BenchmarkScorePositionSettingSourceDate
LMArena Text1125#266 of 297, top 90%LMArena2026-10-08
LMArena Text1094LMArena2026-10-08
LMArena Creative Writing1092#266 of 295, top 91%LMArena2026-10-08
LMArena Creative Writing1035LMArena2026-10-08
EQ-Bench Creative Writing451#114 of 115, top 100%EQ-Bench
LMArena Multi-Turn1117#259 of 295, top 88%LMArena2026-10-08
LMArena Multi-Turn1078LMArena2026-10-08

API pricing by provider

GPT-3.5-turbo API prices
RouteInput $/MOutput $/MCached input $/MChecked
azure$0.50$1.50—2026-10-10
openai$0.50$1.50Free2026-10-10
openrouter$0.50$1.50—2026-10-10

Compare GPT-3.5-turbo

Other OpenAI models

Frequently asked questions

How good is GPT-3.5-turbo?

GPT-3.5-turbo by OpenAI ranks 350th of 354 ranked models on the Noometry Index as of October 2026, with a score of 23.2. Its strongest category is long context, where it ranks 254th. API pricing starts at $0.50 per million input tokens and $1.50 per million output tokens, with a 16K-token context window.

How much does GPT-3.5-turbo cost?

GPT-3.5-turbo costs $0.50 per million input tokens and $1.50 per million output tokens on OpenAI's own API.

What is GPT-3.5-turbo's context window?

GPT-3.5-turbo accepts up to 16K tokens of input and can write up to 4K tokens in one response.

Is GPT-3.5-turbo open source?

No. GPT-3.5-turbo is proprietary and available only through OpenAI's API and partner platforms.

What are GPT-3.5-turbo's strengths and weaknesses?

Relative to other ranked models, GPT-3.5-turbo places best in long context, instruction following, multilingual and lowest in math, writing & preference, coding.

What is GPT-3.5-turbo best at?

Its best category is long context, where it ranks 254th on Noometry.