Terminal-Bench 4.0
AI benchmark · Coding
- Publisher
- Anthropic Fable Mythos 5 1 Launch
- Snapshot

About Terminal-Bench 4.0
Terminal-Bench 4.0 evaluates AI agents on practical terminal-based tasks in isolated environments. It measures the percentage of tasks an agent resolves successfully under a defined model and harness configuration.
Where Terminal-Bench 4.0 sits
Its domains, and the nearest entries sharing them. Click any node to open its page.
The map could not load. The domain links still work.
Next: browse all benchmarks. Entry from the RL Research daily scrape of public sources, 2026-10-10 snapshot.