Long-Horizon Terminal-Bench
AI benchmark · Agentic
- Publisher
- Paper Arxiv
- Domains
- Agentic
- Data source
- benchmarklist
- Snapshot
About Long-Horizon Terminal-Bench
Long-Horizon Terminal-Bench evaluates whether agents can sustain progress on 46 reproducible terminal tasks spanning nine categories, including software engineering, scientific computing, multimodal analysis, research reproduction, systems, professional workflows, games, and logic puzzles. Tasks run in containers under a 90-minute budget and use hidden replay-based verifiers with continuous partial credit, stressing long-horizon planning, context management, iterative debugging, and recovery across hundreds of dependent actions.
Where Long-Horizon Terminal-Bench sits
Its domains, and the nearest entries sharing them. Click any node to open its page.
Select a node to trace its connections. Zoom in for more space; drag to pan.
Next: browse all benchmarks. Entry from the RL Research daily scrape of public sources, 2026-09-21 snapshot.