
Laya vs Jev: Did Decision AI Arrive a Year Earlier?
Laya's creator published related research before Jev, but the public evidence establishes a precursor rather than an identical system.
Original data and analysis on the RL environment economy: vendors, benchmarks, environment craft, and the infrastructure underneath, from the RL Research team.
Meet the authors →
Laya's creator published related research before Jev, but the public evidence establishes a precursor rather than an identical system.

Jev can select bounded actions in RL environments; tasks that require writing still need a generative model.

Automation has reached the loop's last human stage: expert researchers now serve as the measured baseline, and the human role narrows to defining failures and judging what counts as better.

Broad expert task portfolios, not benchmark targeting, are now the reliable training recipe: one 1,700-task set lifted five external coding benchmarks at once and made the agent faster as well as stronger.

When a lab can point 10,000 agents at one problem, unpublished expert work becomes training input, and provenance becomes the priced feature of expert data.

The self-improvement race the resignation describes runs on a supply chain of RL environments and verifiers, and human grading is the only stage of the loop that does not scale with compute.

When agents learn to game a benchmark, the fix is recalibrating the judge, and mature vendors now publish what that recalibration costs.

Expert-built RL environments moved a 397B open-weight model by 70 percent relative on held-out professional tasks, and the recipe is published rather than proprietary.

Composite AI leaderboard scores behave like instruments that get recalibrated, so most single-day movement says more about the ruler than about the models.

Human experts define ground truth for AI agents: authors create private truth, reviewers adjudicate it, and models never approve it. Trusted environments publish their calibration thresholds.

Long-horizon episodes need the operating system and applications the work runs on, plus hard isolation; the pattern the reference spec adopts is microVMs with zero inbound ports.

Under RL pressure agents game environments in predictable ways, from format masquerading as competence to memorization of leaked test structure. The counters are integrity gates that run first and carry zero reward weight.

Environment training transfers across domains. A model post-trained on office tasks with no coding in the mix gained 5.8 points on SWE-Bench Pro, which changes what labs are buying.

Mercor and Ramp's APEX-Accounting is a benchmark for month-end close and bookkeeping. Audit is outside its scope by Mercor's own note, and audit work is graded on evidence integrity, which no reconciliation benchmark carries.

Tracked environments rarely publish verifier calibration or a frozen corpus; a production-grade spec fixes state machine, reward, and splits in a contract before any code.

Jobs whose core work happens on a computer and produces a checkable output can become RL environments; jobs with a critical physical dependency cannot, at any budget.

Mercor's Deeptune acquisition says the constraint has shifted from expert networks to the environments themselves. The vendor data shows the gap the deal targets.

The benchmark pile keeps growing while cataloged RL environments are rare: 99 against 3,029 benchmarks. Expert verification, not benchmark count, moves the numbers.

Joining 202 O*NET occupational tasks against 223 benchmarks with an LLM judge yields definitive coverage of 3.5%. Benchmark success overstates workflow competence.

Marketplace listings are a leading indicator: labs recruit experts months before the environments those experts build reach training runs.

A cottage industry of mostly sub-50-person vendors supplies the most capitalized labs on earth.

Frontier labs are shifting spend from labeled data to executable environments, and vendor margins show the leverage sits in verification rather than labor.