Knowledge benchmark
MMLU leaderboard
As of October 2026, GPT-4o has the highest published MMLU score on Noometry at 88.1%, out of 81 models with results.
Last verified
About MMLU
57-subject multiple-choice exam covering STEM, humanities and professional topics.
Top 15 models
- GPT-4o 88.1%
- Claude 3.5 Sonnet 87.3%
- DeepSeek-V3 87.2%
- Gemini 1.5 Pro (May 2024) 86.9%
- GPT-4 86.4%
- Llama-3.3-70B-Instruct 86.3%
- Qwen2.5 72B Instruct 85.3%
- Phi-4 84.8%
- Claude 3 Opus 84.6%
- Llama 3.1-405B 84.5%
- Qwen2-72B 82.4%
- Amazon Nova Pro 82%
- GPT-4o mini 81.8%
- GPT-4 Turbo 81.3%
- Llama 3.2 90B 80.3%
Sponsored placements are available on pages like this one. Advertise on Noometry
All results
Compare the leaders
Other knowledge benchmarks
- GPQA Diamond
- Humanity's Last Exam
- SimpleQA Verified
- MMLU-Pro
- Confabulations
- Vectara Hallucination Rate
- LMArena Expert
- GPQA (HELM)
- ARC (AI2) Challenge (reference)
- BoolQ (reference)
- OpenBookQA (reference)
- TriviaQA (reference)
Frequently asked questions
What does MMLU measure?
57-subject multiple-choice exam covering STEM, humanities and professional topics.
Which model has the highest MMLU score?
As of October 2026, GPT-4o has the highest published MMLU score on Noometry at 88.1%, out of 81 models with results.
What is the best open-weight model on MMLU?
DeepSeek-V3 has the highest MMLU accuracy among open-weight models at 87.2%, ranking 3 of 81 overall.