Terminal-Bench 4.0

AI benchmark · Coding

Publisher
Anthropic Fable Mythos 5 1 Launch
Snapshot

About Terminal-Bench 4.0

Terminal-Bench 4.0 evaluates AI agents on practical terminal-based tasks in isolated environments. It measures the percentage of tasks an agent resolves successfully under a defined model and harness configuration.

Where Terminal-Bench 4.0 sits

Its domains, and the nearest entries sharing them. Click any node to open its page.

Next: browse all benchmarks. Entry from the RL Research daily scrape of public sources, 2026-10-10 snapshot.