craft 6 min read

Which benchmark version did the agent score on? Harbor's rerun, regrade, reuse rule

Terminal-Bench fixed 28 of 89 tasks in May 2026, then removed 8 and fixed 19 more in August

Harbor’s task version has three digits: environment changes rerun, verifier changes regrade, metadata changes reuse

Agent benchmarks now change faster than the models they grade, and Harbor, the open evaluation harness from the Terminal-Bench team at Laude Institute, answers with a version number on every task. The scheme has three axes. A change to the agent’s environment bumps the first digit and forces a rerun; a change to the verifier, the program that grades each attempt, bumps the second and triggers a regrade of saved trajectories; a metadata change bumps the third and lets old results be reused. The need is measurable: Terminal-Bench 2.1 fixed 28 of 89 tasks on May 6, 2026 and one agent’s score moved 12.1%, then Terminal-Bench 4.0 removed 8 tasks and fixed 19 more on August 28. Ryan Marten of Harbor presented the rule in the Evals track at Runtime, Modal’s conference in San Francisco, on October 1, 2026. A leaderboard without the second axis cannot say whether a score moved because the model changed or because the grader did.

Key Takeaways

  • Harbor versions each task as (x, y, z): x for environment changes (rerun), y for verifier changes (regrade), z for metadata (reuse). Same version and digest, or a patch-only difference, means saved results stand.
  • Terminal-Bench went from 2.0 to 4.0 between May 6 and August 28, 2026: 28 tasks fixed, then a new 74-task set across 7 domains, then 8 tasks removed and 19 fixed under a flat 8-hour agent timeout.
  • A regrade re-scores saved agent outputs with the new verifier and spends no rollout tokens, which makes verifier fixes cheap to apply and cheap to disclose.

What do rerun, regrade, and reuse mean?

Harbor’s rule, written up by Marten in a July 30, 2026 post under the heading “Benchmarks are software and should be maintained like software”, gives every task a three-part version and ties each part to what a maintainer must do with old results.

  • Rerun, (x+1, y, z). The agent environment changed: a new task, a fix to an underspecified or overspecified prompt, modified data or tools, or adjusted agent resources. Old trials are invalid; agents run again.
  • Regrade, (x, y+1, z). The verifier changed: an overly strict check loosened, a reward hack closed, a judge re-tuned, flakiness reduced, verifier resources adjusted, extra metrics reported, or the reward itself changed. Saved trajectories are scored again by the new verifier.
  • Reuse, (x, y, z+1). Metadata changed: documentation, a typo in the prompt, a pinned dependency, oracle flakiness. Old results stand.
Harbor's versioning slide: rerun for agent environment changes, regrade for verifier changes, reuse for metadata changes

Photo: rlresearch.ai, Runtime conference, The Midway, San Francisco, October 1, 2026

Marten's versioning slide at Runtime. The middle column is the one most leaderboards lack: a separate version for the grader.

When a leaderboard moves to a new dataset version, Harbor compares each task’s version and digest and picks one of the three actions per task, so a release that touches ten tasks reruns ten rather than the whole suite. The payoff Marten names for the middle axis: “We can simply re-grade the saved artifacts from previous experiments without burning tokens on new rollouts.”

How much has Terminal-Bench changed in five months?

Terminal-Bench is Harbor’s flagship suite of terminal tasks with automatically checkable success criteria, and its own history shows why the rule exists.

Terminal-Bench versions, May to August 2026: 2.1 fixed 28 of 89 tasks, 3.0 reset to 74 tasks, 4.0 removed 8 and fixed 19 and required a rerun

  • 2.1, May 6, 2026: 28 of the 89 tasks in 2.0 fixed, 9 for external dependencies, 8 for resource mismatches, the rest for misspecification. Most model and agent pairs improved; Claude Code with Opus 4.6 gained 12.1%, and individual task pass rates moved by as much as 84.3% up and 18.6% down, as reported.
  • 3.0, July 30, 2026: a harder set of 74 tasks across 7 domains. Top scores at release: GPT-5.6 Sol in Codex at 34.4%, Fable 5 in Claude Code at 33.8%.
  • 4.0, August 28, 2026: 8 tasks removed (2 saturated, 2 drawing refusals, 2 with public solutions, 2 with quality or platform problems), 19 fixed, and every task set to a flat 8-hour agent timeout. Harbor classed it as a rerun because resources and the task set are environment changes. The release note’s own summary: “TB 4.0 has fewer agent timeouts and errors than 3.0, reducing measurement noise.”

Our own tracking found 135 composite model scores changed in a single refresh on August 29, 2026. A benchmark that fixes 28 tasks in one revision moves its ruler by amounts that swamp most model-to-model gaps.

Why does the verifier get its own version axis?

Because a verifier change alters the score of every past attempt without any agent doing anything new. Regrade re-scores saved artifacts, and Harbor added a regrade command on pull requests so a maintainer can see how a proposed verifier change moves existing trials before merging it. The minor-change list reads like our catalog of agent cheats: close a reward hack, re-tune a judge, reduce flakiness. When Mercor patched hedging in APEX-Agents 1.1, it published the cost, a judge false-negative rate rising from 5.3% to 8.0%; Harbor turns that kind of disclosure into a version digit.

Harbor's framing slide: traditional software treats code as the source of truth; agentic software treats evals as the source of truth

Photo: rlresearch.ai, Runtime conference, The Midway, San Francisco, October 1, 2026

Harbor's opening slide. If evals are the source of truth for learned software, the grader's provenance matters as much as the code's.

Harbor’s opening slide framed evals rather than code as the source of truth for software that is learned. On that framing a verifier edit is an edit to the truth, which is the reason trusted environments publish their calibration.

What this means

An environment that ships without a version on its verifier cannot be compared with itself a month later. The production-grade spec already requires byte-identical replay from seed on the task side; a three-part version on every reported score is the matching requirement on the grading side.

FAQ

What is Harbor?

Harbor is an open framework from Laude Institute, the creators of Terminal-Bench, for specifying sandboxed agent tasks and running evaluations at scale. It installs as a command-line tool, hosts the Terminal-Bench leaderboards, and now carries the task versioning described here.

Is a verifier fix a new benchmark?

Under Harbor’s rule it is a minor version of the same benchmark. Saved agent outputs are regraded, scores on the leaderboard change, and the version digit records why. A new task set or a change to agent resources is a major version and requires rerunning agents.

Why did Terminal-Bench 4.0 need a rerun rather than a regrade?

Because it changed the agent environment: every task moved to a flat 8-hour timeout and 8 tasks were removed. Both are first-digit changes, so the 19 task fixes rode along in a release that had to rerun agents anyway.