TAU-bench
AI benchmark · Agentic
- Domains
- Agentic
- Data source
- benchlm-benchmarks
- Snapshot
About TAU-bench
Original TAU-bench evaluates a model-driven agent in simulated airline and retail customer-service conversations with domain tools, database state, and policy constraints.
Where TAU-bench sits
Its domains, and the nearest entries sharing them. Click any node to open its page.
Select a node to trace its connections. Zoom in for more space; drag to pan.
Next: browse all benchmarks. Entry from the RL Research daily scrape of public sources, 2026-09-21 snapshot.