Agentic & Tool Use benchmark
τ²-bench Telecom leaderboard
As of October 2026, Qwen3.5 397B-A17B has the highest published τ²-bench Telecom score on Noometry at 97.8%, out of 7 models with results.
Last verified
About τ²-bench Telecom
Telecom troubleshooting where both the agent and the simulated user have tools, so the agent has to guide the user through actions on their own device.
- Category
- Agentic & Tool Use
- Introduced
- 2025
- Format
- Tool use with a simulated user
- Unit
- Percent (random guessing ≈ 0%)
- Official site
- taubench.com
Top 7 models
- Qwen3.5 397B-A17B 97.8%
- Claude Opus 4.5 92.3%
- Gemini 3 Flash Preview 91.2%
- Gemini 3 Pro 91%
- GPT-5.2 89.7%
- GLM-5 86.8%
- Claude Sonnet 4.5 84.9%
Sponsored placements are available on pages like this one. Advertise on Noometry
All results
| # | Model | Provider | Score | Setting | Source | Date |
|---|---|---|---|---|---|---|
| 1 | Qwen3.5 397B-A17B | 97.8% | enabled | τ²-bench | 2026-03-02 | |
| 2 | Claude Opus 4.5 | Anthropic | 92.3% | high | τ²-bench | 2026-02-26 |
| 3 | Gemini 3 Flash Preview | 91.2% | high | τ²-bench | 2026-03-02 | |
| 4 | Gemini 3 Pro | 91% | high | τ²-bench | 2026-03-02 | |
| 5 | GPT-5.2 | OpenAI | 89.7% | high | τ²-bench | 2026-02-26 |
| 6 | GLM-5 | Z.ai (Zhipu) | 86.8% | enabled | τ²-bench | 2026-03-02 |
| 7 | Claude Sonnet 4.5 | Anthropic | 84.9% | enabled | τ²-bench | 2026-02-26 |
Compare the leaders
Frequently asked questions
What does τ²-bench Telecom measure?
Telecom troubleshooting where both the agent and the simulated user have tools, so the agent has to guide the user through actions on their own device.
Which model has the highest τ²-bench Telecom score?
As of October 2026, Qwen3.5 397B-A17B has the highest published τ²-bench Telecom score on Noometry at 97.8%, out of 7 models with results.
What is the best open-weight model on τ²-bench Telecom?
Qwen3.5 397B-A17B has the highest τ²-bench Telecom accuracy among open-weight models at 97.8%, ranking 1 of 7 overall.