Agentic & Tool Use benchmark
τ²-bench Airline leaderboard
As of October 2026, Claude Opus 4.5 has the highest published τ²-bench Airline score on Noometry at 84%, out of 7 models with results.
Last verified
About τ²-bench Airline
Airline customer-service conversations where the agent must use booking tools and follow a written policy while a simulated customer makes requests.
- Category
- Agentic & Tool Use
- Introduced
- 2025
- Format
- Tool use with a simulated user
- Unit
- Percent (random guessing ≈ 0%)
- Official site
- taubench.com
Top 7 models
- Claude Opus 4.5 84%
- GPT-5.2 83%
- Gemini 3 Flash Preview 82.5%
- GLM-5 82.5%
- Qwen3.5 397B-A17B 81.5%
- Gemini 3 Pro 80.5%
- Claude Sonnet 4.5 72%
Sponsored placements are available on pages like this one. Advertise on Noometry
All results
| # | Model | Provider | Score | Setting | Source | Date |
|---|---|---|---|---|---|---|
| 1 | Claude Opus 4.5 | Anthropic | 84% | high | τ²-bench | 2026-02-26 |
| 2 | GPT-5.2 | OpenAI | 83% | high | τ²-bench | 2026-02-26 |
| 3 | Gemini 3 Flash Preview | 82.5% | high | τ²-bench | 2026-03-02 | |
| 4 | GLM-5 | Z.ai (Zhipu) | 82.5% | enabled | τ²-bench | 2026-03-02 |
| 5 | Qwen3.5 397B-A17B | 81.5% | enabled | τ²-bench | 2026-03-02 | |
| 6 | Gemini 3 Pro | 80.5% | high | τ²-bench | 2026-03-02 | |
| 7 | Claude Sonnet 4.5 | Anthropic | 72% | enabled | τ²-bench | 2026-02-26 |
Compare the leaders
Frequently asked questions
What does τ²-bench Airline measure?
Airline customer-service conversations where the agent must use booking tools and follow a written policy while a simulated customer makes requests.
Which model has the highest τ²-bench Airline score?
As of October 2026, Claude Opus 4.5 has the highest published τ²-bench Airline score on Noometry at 84%, out of 7 models with results.
What is the best open-weight model on τ²-bench Airline?
GLM-5 has the highest τ²-bench Airline accuracy among open-weight models at 82.5%, ranking 4 of 7 overall.