Knowledge benchmark
BoolQ leaderboard
As of October 2026, Falcon-180B has the highest published BoolQ score on Noometry at 89%, out of 23 models with results.
Last verified
About BoolQ
Naturally occurring yes/no questions about a passage.
- Category
- Knowledge
- Introduced
- 2019
- Format
- Yes/no
- Unit
- Percent (random guessing ≈ 50%)
- Official site
- github.com
Top 15 models
- Falcon-180B 89%
- GPT-4o mini 88.7%
- Llama 2-70B 88.6%
- Mistral 7B 87.4%
- GPT-3.5-turbo 87%
- Qwen-14B 86.2%
- Gemini 1.5 Flash (May 2024) 85.8%
- Gemma 2 9B 85.7%
- Phi-3.5-MoE 84.6%
- Llama 2-34B 83.7%
- Gemma 7B 83.2%
- Falcon-40B 83.1%
- Llama 3.1-8B 82.8%
- Mistral Nemo 82.5%
- Llama 2-13B 82.4%
Sponsored placements are available on pages like this one. Advertise on Noometry
All results
| # | Model | Provider | Score | Setting | Source | Date |
|---|---|---|---|---|---|---|
| 1 | Falcon-180B | 89% | Epoch AI | |||
| 2 | GPT-4o mini | OpenAI | 88.7% | Epoch AI | ||
| 3 | Llama 2-70B | 88.6% | Epoch AI | |||
| 4 | Mistral 7B | 87.4% | Epoch AI | |||
| 5 | GPT-3.5-turbo | OpenAI | 87% | Epoch AI | ||
| 6 | Qwen-14B | 86.2% | Epoch AI | |||
| 7 | Gemini 1.5 Flash (May 2024) | 85.8% | Epoch AI | |||
| 8 | Gemma 2 9B | 85.7% | Epoch AI | |||
| 9 | Phi-3.5-MoE | 84.6% | Epoch AI | |||
| 10 | Llama 2-34B | 83.7% | Epoch AI | |||
| 11 | Gemma 7B | 83.2% | Epoch AI | |||
| 12 | Falcon-40B | 83.1% | Epoch AI | |||
| 13 | Llama 3.1-8B | 82.8% | Epoch AI | |||
| 14 | Mistral Nemo | 82.5% | Epoch AI | |||
| 15 | Llama 2-13B | 82.4% | Epoch AI | |||
| 16 | Llama 13b | 78.7% | Epoch AI | |||
| 17 | Phi-3.5-mini | 78% | Epoch AI | |||
| 18 | Llama 2-7B | 77.9% | Epoch AI | |||
| 19 | Qwen-7B | 76.4% | Epoch AI | |||
| 20 | Phi-1.5 | 75.8% | 5 | Epoch AI | ||
| 21 | Falcon-7B | 75.3% | Epoch AI | |||
| 22 | Gemma 2B | 69.4% | Epoch AI | |||
| 23 | Dolly 2.0-12b | 56.3% | Epoch AI |
Compare the leaders
Other knowledge benchmarks
- GPQA Diamond
- Humanity's Last Exam
- SimpleQA Verified
- MMLU-Pro
- Confabulations
- Vectara Hallucination Rate
- LMArena Expert
- GPQA (HELM)
- ARC (AI2) Challenge (reference)
- MMLU (reference)
- OpenBookQA (reference)
- TriviaQA (reference)
Frequently asked questions
What does BoolQ measure?
Naturally occurring yes/no questions about a passage.
Which model has the highest BoolQ score?
As of October 2026, Falcon-180B has the highest published BoolQ score on Noometry at 89%, out of 23 models with results.
What is the best open-weight model on BoolQ?
Falcon-180B has the highest BoolQ accuracy among open-weight models at 89%, ranking 1 of 23 overall.