Which jobs can become RL environments?
By Elon HuskParody pen name · View profile10 screening dimensions, 1 binary veto (critical physical dependency)

Jobs whose core work happens on a computer and produces a checkable output can become RL environments; jobs with a critical physical dependency cannot, at any budget. rlresearch.ai’s screening framework scores each job on ten dimensions, from computer-work share to output verifiability, with one hard rule: a critical physical dependency vetoes construction outright, with no score adjustment. A job that requires hands-on equipment operation, physical inventory observation, or in-person client interaction cannot become a sandbox environment, however well it scores everywhere else. Three reference workflows screened with the framework, failure diagnosis, engineering test analysis, and audit workpaper completion, pass the gate with computer-work shares of 0.82 to 0.96, and each names the physical capabilities it excludes. Surge AI’s Chartography, launched July 16, 2026, passes every screen. The framework sorts the 66 tracked occupations without environments into vetoed and not yet built.
Key Takeaways
- Ten dimensions decide whether a job can become an RL environment, and nine of them are scored. A critical physical dependency is a binary veto.
- Three screened workflows pass the gate with computer-work shares of 0.82 to 0.96, and each names which physical capabilities it excludes.
- The framework predicts the market: Surge AI’s Chartography (July 16, 2026) passes every screen, while construction trades and farming stay unbuildable at any budget.
What are the ten screening dimensions?
Each candidate occupation is scored on nine dimensions from 0 to 1. Five describe the shape of the work:
- workflow modality
- computer-work share
- sandbox-action coverage
- native-tool fidelity
- long-horizon depth
Four describe whether the work can be observed and graded:
- digital-evidence availability
- action observability
- output verifiability
- collaboration simulability
The tenth, critical physical dependency, is a binary veto and decides eligibility on its own. A job whose core competency requires physical manipulation of equipment, hands-on bench operation, or in-person client interaction cannot be simulated at a desk, and no amount of computer-work share compensates. The framework rejects the job before scoring the other dimensions, because a sandbox for a physical job measures tool adaptation rather than competence.
OpenAI’s GPTs are GPTs scored ONET tasks (from the US Department of Labor’s occupational task database) for LLM exposure and excluded the same tasks by rubric definition: tasks requiring physical action scored zero exposure. Anthropic’s mapping of millions of real conversations onto ONET tasks shows usage concentrating in the computer-mediated task families the framework passes (Handa et al., 2025).
What do three screened workflows show?
Three reference workflow specs screened by rlresearch.ai illustrate the framework:
- Failure diagnosis (remote reliability triage and maintenance work-package authoring) scores 0.82 on computer-work share. The scoped workflow excludes physical inspection, lockout execution, and repair workmanship. The desk-based triage is a separate competency.
- Engineering test analysis scores 0.92. The scoped workflow excludes hands-on bench operation and sensor installation. The data audit, design critique, and decision-making are computer-mediated.
- Audit workpaper completion scores 0.96. Audit is computer-native: source records, formulas, evidence links, and conclusions are all digital. What its spec contains is covered separately.
All three pass the screening gate. None has a critical physical dependency for the scoped workflow, and each names the capabilities it excludes.
Two published benchmarks show the pass side of the gate. SWE-Lancer converted 1,400+ real freelance software jobs, with real dollar values attached, into end-to-end verifiable tasks. GDPval built expert-graded tasks from real work products across 44 occupations.
Does Surge AI’s Chartography pass the screen?
Yes. Surge AI launched Chartography on July 16, 2026: expert-graded professional chart reading. A chart-reading task requires no physical manipulation, produces a digital artifact, and can be graded against expert rubrics. It is the kind of work the framework predicts is environment-eligible.
The 66 tracked occupations without environments include construction trades and farming, jobs with critical physical dependencies that veto sandbox construction. They also include legal work and management consulting, jobs that score high on computer-work share and output verifiability but have no environment yet. The framework separates the two groups.
“introducing CHARTOGRAPHY, our new benchmark for professional chart understanding”
Chart reading is fully computer-mediated and its output is checkable, which is why it passes every screen in the framework.
What this means
The screening framework is the first filter in rlresearch.ai’s reference designs. Occupations that pass are buildable; occupations vetoed by physical dependency are not. Among the passing occupations, the ones with no environment yet are where recruiting demand already points.
FAQ
What is a critical physical dependency?
A job has a critical physical dependency when its core competency requires manipulating physical objects (equipment, inventory, patients, materials) in ways that cannot be represented at a desk. The framework vetoes these jobs outright because a sandbox simulation would measure adaptation to substitute tools instead of the competence the job requires.
Why is native-tool fidelity a measured variable?
Native-tool fidelity is scored because grading workbook formulas, review behavior, and final artifacts depends on the native tool; the score records how much the tool choice changes what is measured. The reference audit spec runs in Windows with native Office for that reason.
What is Chartography?
Chartography is a Surge AI benchmark launched July 16, 2026 for expert-graded professional chart reading. It passes every screen in the framework: fully computer-mediated, deterministic evidence, verifiable output.
