Agentic & Tool Use benchmark

τ²-bench Retail leaderboard

As of October 2026, Qwen3.5 397B-A17B has the highest published τ²-bench Retail score on Noometry at 84.4%, out of 7 models with results.

Last verified

About τ²-bench Retail

Retail customer-service conversations (returns, exchanges, order changes) handled with tools under a written policy.

Category
Agentic & Tool Use
Introduced
2025
Format
Tool use with a simulated user
Unit
Percent (random guessing ≈ 0%)
Official site
taubench.com

Top 7 models

Top models on τ²-bench Retail
  1. Qwen3.5 397B-A17B 84.4%
  2. GPT-5.2 81.6%
  3. Claude Opus 4.5 79.6%
  4. Gemini 3 Flash Preview 76.8%
  5. Gemini 3 Pro 75.9%
  6. GLM-5 73.7%
  7. Claude Sonnet 4.5 72.4%

Sponsored placements are available on pages like this one. Advertise on Noometry

All results

τ²-bench Retail results by model
#ModelProviderScoreSettingSourceDate
1Qwen3.5 397B-A17B Alibaba (Qwen)84.4%enabledτ²-bench2026-03-02
2GPT-5.2 OpenAI81.6%highτ²-bench2026-02-26
3Claude Opus 4.5 Anthropic79.6%highτ²-bench2026-02-26
4Gemini 3 Flash Preview Google76.8%highτ²-bench2026-03-02
5Gemini 3 Pro Google75.9%highτ²-bench2026-03-02
6GLM-5 Z.ai (Zhipu)73.7%enabledτ²-bench2026-03-02
7Claude Sonnet 4.5 Anthropic72.4%enabledτ²-bench2026-02-26

Compare the leaders

Other agentic & tool use benchmarks

Frequently asked questions

What does τ²-bench Retail measure?

Retail customer-service conversations (returns, exchanges, order changes) handled with tools under a written policy.

Which model has the highest τ²-bench Retail score?

As of October 2026, Qwen3.5 397B-A17B has the highest published τ²-bench Retail score on Noometry at 84.4%, out of 7 models with results.

What is the best open-weight model on τ²-bench Retail?

Qwen3.5 397B-A17B has the highest τ²-bench Retail accuracy among open-weight models at 84.4%, ranking 1 of 7 overall.