Long-Horizon Terminal-Bench

AI benchmark · Agentic

Publisher
Paper Arxiv
Domains
Agentic
Data source
benchmarklist
Snapshot

Open benchmark ↗benchmarklist.com/benchmarks/long_horizon_terminal_bench/

About Long-Horizon Terminal-Bench

Long-Horizon Terminal-Bench evaluates whether agents can sustain progress on 46 reproducible terminal tasks spanning nine categories, including software engineering, scientific computing, multimodal analysis, research reproduction, systems, professional workflows, games, and logic puzzles. Tasks run in containers under a 90-minute budget and use hidden replay-based verifiers with continuous partial credit, stressing long-horizon planning, context management, iterative debugging, and recovery across hundreds of dependent actions.

Where Long-Horizon Terminal-Bench sits

Its domains, and the nearest entries sharing them. Click any node to open its page.

Select a node to trace its connections. Zoom in for more space; drag to pan.

Next: browse all benchmarks. Entry from the RL Research daily scrape of public sources, 2026-09-21 snapshot.