Agentic & Tool Use benchmark
Berkeley Function Calling Leaderboard leaderboard
As of October 2026, Claude Opus 4.5 has the highest published Berkeley Function Calling Leaderboard score on Noometry at 77.5%, out of 49 models with results.
Last verified
About Berkeley Function Calling Leaderboard
Tool-calling accuracy across single, parallel and multi-turn function calls, web search, memory and knowing when not to call a tool (BFCL v4).
- Category
- Agentic & Tool Use
- Introduced
- 2024
- Format
- Function calls
- Unit
- Percent (random guessing ≈ 0%)
- Official site
- gorilla.cs.berkeley.edu
Top 15 models
- Claude Opus 4.5 77.5%
- Claude Sonnet 4.5 73.2%
- Gemini 3 Pro 72.5%
- GLM-4.6 72.4%
- Grok 4.1 Fast 69.6%
- Claude Haiku 4.5 68.7%
- o3 63%
- Grok 4 63%
- Kimi K2 (Jul 2025) 59.1%
- Command A 57.1%
- DeepSeek-V3.2-Exp 56.7%
- Gemini 2.5 Flash 56.2%
- GPT-5.2 55.9%
- GPT-5 Mini 55.5%
- GPT-4.1 54%
Sponsored placements are available on pages like this one. Advertise on Noometry
All results
Compare the leaders
Frequently asked questions
What does Berkeley Function Calling Leaderboard measure?
Tool-calling accuracy across single, parallel and multi-turn function calls, web search, memory and knowing when not to call a tool (BFCL v4).
Which model has the highest Berkeley Function Calling Leaderboard score?
As of October 2026, Claude Opus 4.5 has the highest published Berkeley Function Calling Leaderboard score on Noometry at 77.5%, out of 49 models with results.
What is the best open-weight model on Berkeley Function Calling Leaderboard?
GLM-4.6 has the highest Berkeley Function Calling Leaderboard accuracy among open-weight models at 72.4%, ranking 4 of 49 overall.