craft 5 min read

Four vendors published RL environment checklists. They agree on resets and exploits, and split on partial credit

4 published checklists; 3 shared requirements; 1 split (partial credit: 27.25 checkpoints per OSWorld 2.0 workflow, 20.6% end-to-end vs 54.8% partial); 0 publish calibration thresholds

Four vendor checklists agree on three environment requirements, split on partial credit, and none publishes calibration thresholds

Four vendors have now published, in their own words, what a reinforcement learning (RL) environment must have before anyone trains or evaluates on it, and the lists agree more than their marketing does. Snorkel’s design guide of September 28, 2026, Scale’s launch post from February, Surge’s grading rules for HANDBOOK.md and DAYJOB, and Harbor’s versioning rule converge on three requirements: a deterministic reset, an explicit test for reward exploits, and a versioned harness. All three are in the reference spec rlresearch.ai published in July. The split is partial credit. Snorkel recommends checkpoint scoring for long workflows, citing OSWorld 2.0’s 27.25 checkpoints per workflow and a top system at 20.6% end-to-end but 54.8% partial; Surge grades a task as passed only when every criterion is met; the reference spec pays graded credit only for wrong-but-valid answers. None of the four publishes verifier calibration thresholds, which is the item the spec adds.

Key Takeaways

  • Shared by all four: a known starting state that resets reliably and is confirmed solvable, an adversarial test of whether the verifier can be satisfied without doing the task, and versioning of the harness and verifier so repeated runs test the same thing.
  • The split: Snorkel’s partial credit for long workflows against Surge’s all-criteria pass and the reference spec’s credit only for wrong-but-valid answers.
  • Missing from all four: published precision and recall for the verifier against expert adjudication, which the reference spec sets at 0.95 or better and at a hard 1.00 for integrity gates.

What does each vendor require?

Snorkel’s guide, by Aryan Kargwal and Jonathan Schlosser, defines the environment as a six-part loop: state, observations, actions, transitions, rewards, and resets, with the verifier at the center. “A verifier is the component that inspects the outcome and decides whether the task was completed.” Its validation checklist has three headings: solvability, confirmed by an oracle run on a fresh reset; reward exploits, tested with adversarial runs in which an agent is told to maximize score by any means, plus a permissions audit of grader code and CI; and drift, which means versioning the harness with the environment and logging which tools were available.

Scale’s February post, by Chetan Rane, lists expert-curated realism, verifiers that check real changes in system state, a defined state that resets between runs, repeatable starting configurations with replayable trajectories, and parallel infrastructure. “Results are evaluated using expert-designed verifiers that check for real, measurable changes in the system”. Surge publishes a grading rule rather than a checklist: HANDBOOK.md grades deterministically against per-task required outputs and prohibited actions, and DAYJOB counts a pass only when every rubric criterion is met. Harbor’s rule versions the agent environment, the verifier, and the metadata separately, so that closing a reward hack or re-tuning a judge is a recorded change that regrades saved trajectories.

Four vendor checklists and the reference spec across five requirements: reset, exploit test, versioning, partial credit, and published calibration

Where do they agree with the reference spec?

On three points. The reset requirement is the spec’s byte-identical replay from seed, and Snorkel’s insistence on a fresh reset still yielding a solvable task is the per-episode reset our infrastructure analysis listed as unsolved at the application layer. The exploit test is the spec’s red-team role: models draft and attack, and an environment ships only after the integrity gates hold against them. Snorkel’s phrasing is the cleanest statement of why: “A verifier can be technically correct and still reward the wrong solution.” Versioning is Harbor’s contribution, and Snorkel’s drift section reaches the same place from the other side, listing new tools, app updates, background processes, and cached state as the ways a benchmark stops measuring what it measured last month.

Where do they split, and what is missing?

The split is partial credit, and it is a real disagreement about what the reward should teach. Snorkel’s case rests on long workflows: “A final pass or fail says very little about a long trajectory.” OSWorld 2.0 averages 27.25 checkpoints per workflow, and the top system there completes 20.6% of workflows end to end while earning 54.8% partial credit, so the checkpoints carry most of the training signal. The guide’s safeguard is that checkpoints describe results rather than required action sequences. Surge takes the other side for professional deliverables: a DAYJOB pass requires every criterion. The reference spec sits between them. It scores one verifiable outcome bit after the gates pass, because multi-component rubrics leak reward through format and keywords, and it pays graded credit only to wrong-but-valid answers so the RL signal stays dense enough to learn from.

What none of the four publishes is calibration. Snorkel asks that each semantic criterion yield consistent scores across repeated runs, which is a reliability requirement, and the others are silent. The spec’s bar is precision and recall of 0.95 or better against adjudicated expert scores and a hard 1.00 for the integrity gates, published with the environment, so a buyer can audit the grader rather than trust it.

What this means

The industry’s checklists have converged on the mechanics of a sound environment, which makes resets, exploit tests, and versioning table stakes rather than differentiators. The open questions a buyer should ask a vendor are the two the checklists leave out: how the reward treats partial completion, and what the verifier’s measured precision and recall are.

FAQ

What is partial credit in an RL environment?

A reward that pays for intermediate progress, typically a weighted share of checkpoints reached, rather than only for completing the task. Snorkel’s formula divides the weighted sum of checkpoints met by the total weight. It gives long workflows a dense training signal at the risk of rewarding progress that never reaches a correct outcome.

What is a reward exploit test?

A run in which an agent is instructed to maximize score by any means, combined with an audit of who can write to the test files, grader code, and CI configuration. If score can be earned without completing the task, the verifier fails the test. Snorkel, the reference spec, and Harbor’s regrade axis all treat this as a precondition for training.

Why does calibration matter if the checklist is met?

A verifier can reset cleanly, resist the exploits that were tried, and be versioned, and still be wrong on a share of episodes. Published precision and recall against expert adjudication tell a buyer how large that share is.