Agentic & Tool Use benchmark

τ²-bench Telecom leaderboard

As of October 2026, Qwen3.5 397B-A17B has the highest published τ²-bench Telecom score on Noometry at 97.8%, out of 7 models with results.

Last verified

About τ²-bench Telecom

Telecom troubleshooting where both the agent and the simulated user have tools, so the agent has to guide the user through actions on their own device.

Category
Agentic & Tool Use
Introduced
2025
Format
Tool use with a simulated user
Unit
Percent (random guessing ≈ 0%)
Official site
taubench.com

Top 7 models

Top models on τ²-bench Telecom
  1. Qwen3.5 397B-A17B 97.8%
  2. Claude Opus 4.5 92.3%
  3. Gemini 3 Flash Preview 91.2%
  4. Gemini 3 Pro 91%
  5. GPT-5.2 89.7%
  6. GLM-5 86.8%
  7. Claude Sonnet 4.5 84.9%

Sponsored placements are available on pages like this one. Advertise on Noometry

All results

τ²-bench Telecom results by model
#ModelProviderScoreSettingSourceDate
1Qwen3.5 397B-A17B Alibaba (Qwen)97.8%enabledτ²-bench2026-03-02
2Claude Opus 4.5 Anthropic92.3%highτ²-bench2026-02-26
3Gemini 3 Flash Preview Google91.2%highτ²-bench2026-03-02
4Gemini 3 Pro Google91%highτ²-bench2026-03-02
5GPT-5.2 OpenAI89.7%highτ²-bench2026-02-26
6GLM-5 Z.ai (Zhipu)86.8%enabledτ²-bench2026-03-02
7Claude Sonnet 4.5 Anthropic84.9%enabledτ²-bench2026-02-26

Compare the leaders

Other agentic & tool use benchmarks

Frequently asked questions

What does τ²-bench Telecom measure?

Telecom troubleshooting where both the agent and the simulated user have tools, so the agent has to guide the user through actions on their own device.

Which model has the highest τ²-bench Telecom score?

As of October 2026, Qwen3.5 397B-A17B has the highest published τ²-bench Telecom score on Noometry at 97.8%, out of 7 models with results.

What is the best open-weight model on τ²-bench Telecom?

Qwen3.5 397B-A17B has the highest τ²-bench Telecom accuracy among open-weight models at 97.8%, ranking 1 of 7 overall.