Terminal-Bench-Science 0.1

AI benchmark · Science

Publisher
Anthropic Fable Mythos 5 1 Launch
Snapshot

About Terminal-Bench-Science 0.1

Terminal-Bench-Science 0.1 evaluates AI agents on 70 outcome-verifiable scientific workflows performed in containerized terminal environments. Tasks span computational biology, chemistry, physics, earth science, and related technical research work.

Where Terminal-Bench-Science 0.1 sits

Its domains, and the nearest entries sharing them. Click any node to open its page.

Next: browse all benchmarks. Entry from the RL Research daily scrape of public sources, 2026-10-10 snapshot.