Agentic & Tool Use benchmark

Berkeley Function Calling Leaderboard leaderboard

As of October 2026, Claude Opus 4.5 has the highest published Berkeley Function Calling Leaderboard score on Noometry at 77.5%, out of 49 models with results.

Last verified

About Berkeley Function Calling Leaderboard

Tool-calling accuracy across single, parallel and multi-turn function calls, web search, memory and knowing when not to call a tool (BFCL v4).

Category
Agentic & Tool Use
Introduced
2024
Format
Function calls
Unit
Percent (random guessing ≈ 0%)
Official site
gorilla.cs.berkeley.edu

Top 15 models

Top models on Berkeley Function Calling Leaderboard
  1. Claude Opus 4.5 77.5%
  2. Claude Sonnet 4.5 73.2%
  3. Gemini 3 Pro 72.5%
  4. GLM-4.6 72.4%
  5. Grok 4.1 Fast 69.6%
  6. Claude Haiku 4.5 68.7%
  7. o3 63%
  8. Grok 4 63%
  9. Kimi K2 (Jul 2025) 59.1%
  10. Command A 57.1%
  11. DeepSeek-V3.2-Exp 56.7%
  12. Gemini 2.5 Flash 56.2%
  13. GPT-5.2 55.9%
  14. GPT-5 Mini 55.5%
  15. GPT-4.1 54%

Sponsored placements are available on pages like this one. Advertise on Noometry

All results

Berkeley Function Calling Leaderboard results by model
#ModelProviderScoreSettingSourceDate
1Claude Opus 4.5 Anthropic77.5%fcBerkeley Function Calling Leaderboard
2Claude Sonnet 4.5 Anthropic73.2%fcBerkeley Function Calling Leaderboard
3Gemini 3 Pro Google72.5%promptBerkeley Function Calling Leaderboard
4GLM-4.6 Z.ai (Zhipu)72.4%fc thinkingBerkeley Function Calling Leaderboard
5Grok 4.1 Fast xAI69.6%fcBerkeley Function Calling Leaderboard
6Claude Haiku 4.5 Anthropic68.7%fcBerkeley Function Calling Leaderboard
7o3 OpenAI63%promptBerkeley Function Calling Leaderboard
8Grok 4 xAI63%promptBerkeley Function Calling Leaderboard
9Kimi K2 (Jul 2025) Moonshot AI59.1%fcBerkeley Function Calling Leaderboard
10Command A Cohere57.1%fcBerkeley Function Calling Leaderboard
11DeepSeek-V3.2-Exp DeepSeek56.7%prompt + thinkingBerkeley Function Calling Leaderboard
12Gemini 2.5 Flash Google56.2%fcBerkeley Function Calling Leaderboard
13GPT-5.2 OpenAI55.9%fcBerkeley Function Calling Leaderboard
14GPT-5 Mini OpenAI55.5%fcBerkeley Function Calling Leaderboard
15GPT-4.1 OpenAI54%fcBerkeley Function Calling Leaderboard
16o4-mini OpenAI53.2%fcBerkeley Function Calling Leaderboard
17Qwen3 235B-A22B Alibaba (Qwen)52.1%promptBerkeley Function Calling Leaderboard
18GPT-5 Nano OpenAI51.5%fcBerkeley Function Calling Leaderboard
19GPT-4.1 mini OpenAI50.5%fcBerkeley Function Calling Leaderboard
20Qwen3 32B Alibaba (Qwen)48.7%fcBerkeley Function Calling Leaderboard
21Qwen3 8B Alibaba (Qwen)42.6%fcBerkeley Function Calling Leaderboard
22Qwen3-30B-A3B Alibaba (Qwen)41.4%fcBerkeley Function Calling Leaderboard
23Qwen3 14B Alibaba (Qwen)41%fcBerkeley Function Calling Leaderboard
24Mistral Large Mistral AI38.4%fcBerkeley Function Calling Leaderboard
25Mistral Medium Mistral AI37.7%Berkeley Function Calling Leaderboard
26Llama 4 Maverick Meta37.3%fcBerkeley Function Calling Leaderboard
27Mistral Small Mistral AI37.1%fcBerkeley Function Calling Leaderboard
28Gemini 2.5 Flash-Lite Google36.9%fcBerkeley Function Calling Leaderboard
29Qwen3-4B Alibaba (Qwen)35.7%fcBerkeley Function Calling Leaderboard
30GPT-4.1 nano OpenAI33%fcBerkeley Function Calling Leaderboard
31Command R7B Cohere32.1%fcBerkeley Function Calling Leaderboard
32Llama-3.3-70B-Instruct Meta31.9%fcBerkeley Function Calling Leaderboard
33Gemma 3 12B Google30.4%promptBerkeley Function Calling Leaderboard
34Gemma 3 27B Google29.5%promptBerkeley Function Calling Leaderboard
35Phi-4 Microsoft28.8%promptBerkeley Function Calling Leaderboard
36Qwen3-1.7B Alibaba (Qwen)28.4%fcBerkeley Function Calling Leaderboard
37Llama 4 Scout Meta28.1%fcBerkeley Function Calling Leaderboard
38Mistral Nemo Mistral AI27.6%fcBerkeley Function Calling Leaderboard
39Granite 3.1 8b Instruct IBM27.1%fcBerkeley Function Calling Leaderboard
40Nova 2 Lite Amazon27.1%fcBerkeley Function Calling Leaderboard
41Llama 3.1-8B Meta25.8%promptBerkeley Function Calling Leaderboard
42Amazon Nova Pro Amazon25%fcBerkeley Function Calling Leaderboard
43Amazon Nova Micro Amazon22.3%fcBerkeley Function Calling Leaderboard
44Llama 3.2 3B Meta21.9%fcBerkeley Function Calling Leaderboard
45Gemma 3 4B Google19.6%promptBerkeley Function Calling Leaderboard
46Ministral 8B Mistral AI11.1%fcBerkeley Function Calling Leaderboard
47Llama 3.2 1B Meta10.8%fcBerkeley Function Calling Leaderboard
48Llama 3.1 Nemotron Ultra 253b v1 NVIDIA10%fcBerkeley Function Calling Leaderboard
49Gemma 3 1B Google7.2%promptBerkeley Function Calling Leaderboard

Compare the leaders

Other agentic & tool use benchmarks

Frequently asked questions

What does Berkeley Function Calling Leaderboard measure?

Tool-calling accuracy across single, parallel and multi-turn function calls, web search, memory and knowing when not to call a tool (BFCL v4).

Which model has the highest Berkeley Function Calling Leaderboard score?

As of October 2026, Claude Opus 4.5 has the highest published Berkeley Function Calling Leaderboard score on Noometry at 77.5%, out of 49 models with results.

What is the best open-weight model on Berkeley Function Calling Leaderboard?

GLM-4.6 has the highest Berkeley Function Calling Leaderboard accuracy among open-weight models at 72.4%, ranking 4 of 49 overall.