Knowledge benchmark
OpenBookQA leaderboard
As of October 2026, Phi 3 Mini 4k Instruct has the highest published OpenBookQA score on Noometry at 88%, out of 19 models with results.
Last verified
About OpenBookQA
Elementary science questions that require combining a given fact with common knowledge.
- Category
- Knowledge
- Introduced
- 2018
- Format
- Multiple choice
- Unit
- Percent (random guessing ≈ 25%)
- Official site
- allenai.org
Top 15 models
- Phi 3 Mini 4k Instruct 88%
- Phi 3 Small 8k Instruct 88%
- phi-3-medium 14B 87.4%
- GPT-3.5-turbo 86%
- Mixtral 8x7B 85.8%
- Llama 3-8B 82.6%
- Mistral 7B 79.8%
- Gemma 7B 78.6%
- Phi-2 73.6%
- Falcon-180B 64.2%
- Llama 2-70B 60.2%
- Llama 2-7B 58.6%
- Llama 2-34B 58.2%
- Llama 2-13B 57%
- Falcon-40B 56.6%
Sponsored placements are available on pages like this one. Advertise on Noometry
All results
| # | Model | Provider | Score | Setting | Source | Date |
|---|---|---|---|---|---|---|
| 1 | Phi 3 Mini 4k Instruct | 88% | Epoch AI | |||
| 2 | Phi 3 Small 8k Instruct | 88% | Epoch AI | |||
| 3 | phi-3-medium 14B | 87.4% | Epoch AI | |||
| 4 | GPT-3.5-turbo | OpenAI | 86% | Epoch AI | ||
| 5 | Mixtral 8x7B | 85.8% | Epoch AI | |||
| 6 | Llama 3-8B | 82.6% | Epoch AI | |||
| 7 | Mistral 7B | 79.8% | Epoch AI | |||
| 8 | Gemma 7B | 78.6% | Epoch AI | |||
| 9 | Phi-2 | 73.6% | Epoch AI | |||
| 10 | Falcon-180B | 64.2% | Epoch AI | |||
| 11 | Llama 2-70B | 60.2% | Epoch AI | |||
| 12 | Llama 2-7B | 58.6% | Epoch AI | |||
| 13 | Llama 2-34B | 58.2% | Epoch AI | |||
| 14 | Llama 2-13B | 57% | Epoch AI | |||
| 15 | Falcon-40B | 56.6% | Epoch AI | |||
| 16 | Llama 13b | 56.4% | Epoch AI | |||
| 17 | Falcon-7B | 51.6% | Epoch AI | |||
| 18 | Dolly 2.0-12b | 39.2% | Epoch AI | |||
| 19 | Phi-1.5 | 37.2% | 5 | Epoch AI |
Compare the leaders
Other knowledge benchmarks
- GPQA Diamond
- Humanity's Last Exam
- SimpleQA Verified
- MMLU-Pro
- Confabulations
- Vectara Hallucination Rate
- LMArena Expert
- GPQA (HELM)
- ARC (AI2) Challenge (reference)
- BoolQ (reference)
- MMLU (reference)
- TriviaQA (reference)
Frequently asked questions
What does OpenBookQA measure?
Elementary science questions that require combining a given fact with common knowledge.
Which model has the highest OpenBookQA score?
As of October 2026, Phi 3 Mini 4k Instruct has the highest published OpenBookQA score on Noometry at 88%, out of 19 models with results.
What is the best open-weight model on OpenBookQA?
Phi 3 Mini 4k Instruct has the highest OpenBookQA accuracy among open-weight models at 88%, ranking 1 of 19 overall.