TAU-bench

AI benchmark · Agentic

Domains
Agentic
Data source
benchlm-benchmarks
Snapshot

Open benchmark ↗benchlm.ai/benchmarks/tauBench

About TAU-bench

Original TAU-bench evaluates a model-driven agent in simulated airline and retail customer-service conversations with domain tools, database state, and policy constraints.

Where TAU-bench sits

Its domains, and the nearest entries sharing them. Click any node to open its page.

Select a node to trace its connections. Zoom in for more space; drag to pan.

Next: browse all benchmarks. Entry from the RL Research daily scrape of public sources, 2026-09-21 snapshot.