Agentic & Tool Use benchmark

τ²-bench Banking leaderboard

As of October 2026, Qwen3.8 Max has the highest published τ²-bench Banking score on Noometry at 55.1%, out of 26 models with results.

Last verified

About τ²-bench Banking

Banking customer-service tasks that need the agent to search a knowledge base as well as use account tools under policy.

Category
Agentic & Tool Use
Format
Tool use with retrieval
Unit
Percent (random guessing ≈ 0%)
Official site
taubench.com

Top 15 models

Top models on τ²-bench Banking
  1. Qwen3.8 Max 55.1%
  2. Claude Opus 5 48.7%
  3. Grok 4.5 47.9%
  4. GPT-5.6 Sol 46.9%
  5. GPT-5.5 44.6%
  6. Muse Spark 1.1 40.5%
  7. Claude Opus 4.7 40.2%
  8. Claude Fable 5 39.7%
  9. Claude Opus 4.8 39.7%
  10. GPT-5.4 39.4%
  11. GLM-5.2 37.1%
  12. Kimi K3 37.1%
  13. GPT-5.2 32.2%
  14. Claude Opus 4.6 27.3%
  15. Gemini 3 Flash Preview 27.3%

Sponsored placements are available on pages like this one. Advertise on Noometry

All results

τ²-bench Banking results by model
#ModelProviderScoreSettingSourceDate
1Qwen3.8 Max Alibaba (Qwen)55.1%xhighτ²-bench2026-08-04
2Claude Opus 5 Anthropic48.7%maxτ²-bench2026-08-04
3Grok 4.5 xAI47.9%highτ²-bench2026-08-04
4GPT-5.6 Sol OpenAI46.9%xhighτ²-bench2026-08-04
5GPT-5.5 OpenAI44.6%xhighτ²-bench2026-05-05
6Muse Spark 1.1 Meta40.5%xhighτ²-bench2026-08-04
7Claude Opus 4.7 Anthropic40.2%maxτ²-bench2026-05-05
8Claude Fable 5 Anthropic39.7%maxτ²-bench2026-08-04
9Claude Opus 4.8 Anthropic39.7%maxτ²-bench2026-08-04
10GPT-5.4 OpenAI39.4%xhighτ²-bench2026-03-25
11GLM-5.2 Z.ai (Zhipu)37.1%xhighτ²-bench2026-08-04
12Kimi K3 Moonshot AI37.1%maxτ²-bench2026-08-04
13GPT-5.2 OpenAI32.2%highτ²-bench2026-02-26
14Claude Opus 4.6 Anthropic27.3%maxτ²-bench2026-05-05
15Gemini 3 Flash Preview Google27.3%highτ²-bench2026-03-02
16Gemini 3.1 Pro Preview Google26%highτ²-bench2026-05-05
17Claude Sonnet 4.5 Anthropic25.3%enabledτ²-bench2026-02-26
18Inkling Thinking Machines Lab25%maxτ²-bench2026-08-04
19Claude Opus 4.5 Anthropic24.7%highτ²-bench2026-02-26
20Gemini 3 Pro Google18%highτ²-bench2026-03-02
21Grok 4.20 (Non-Reasoning) xAI18%highτ²-bench2026-05-05
22Grok 4 Fast xAI15.7%highτ²-bench2026-05-05
23Gemini 2.5 Pro Google13.7%highτ²-bench2026-05-05
24Grok 4.1 Fast xAI13.1%highτ²-bench2026-05-05
25GLM-5 Z.ai (Zhipu)9.8%enabledτ²-bench2026-03-02
26Qwen3.5 397B-A17B Alibaba (Qwen)9.8%enabledτ²-bench2026-03-02

Compare the leaders

Other agentic & tool use benchmarks

Frequently asked questions

What does τ²-bench Banking measure?

Banking customer-service tasks that need the agent to search a knowledge base as well as use account tools under policy.

Which model has the highest τ²-bench Banking score?

As of October 2026, Qwen3.8 Max has the highest published τ²-bench Banking score on Noometry at 55.1%, out of 26 models with results.

What is the best open-weight model on τ²-bench Banking?

GLM-5.2 has the highest τ²-bench Banking accuracy among open-weight models at 37.1%, ranking 11 of 26 overall.