market Updated 4 min read

What do AI expert job postings reveal about RL environment demand?

3 expert marketplaces tracked, cross-posted roles counted once

Expert recruiting on Mercor, AfterQuery, and micro1 spans medical, legal, and finance while 21 of 38 vendors sell coding environments

Expert job listings on Mercor, AfterQuery, and micro1 are a leading indicator of where AI labs are investing in environment development. We track active expert postings across all three platforms, deduplicated so cross-posted roles count once, and classify each by domain using the platform’s own taxonomy first, with title inference only as a fallback. The pattern is consistent: labs recruit domain experts months before the environments those experts build reach a training run. Recruiting concentrates in a few domains (coding, finance, legal, medical), while our vendor census counts 21 of 38 tracked companies selling coding environments, 18 selling enterprise workflows, and 18 selling computer-use environments; math has 2 sellers and medical effectively one. Where recruiting volume diverges from what vendors currently sell is where the next environment categories will appear.

Key Takeaways

  • Labs recruit domain experts on Mercor, AfterQuery, and micro1 months before the resulting environments reach training runs, which makes postings the earliest public demand signal (as of June 2026).
  • Recruiting spans medical, legal, and finance, while 21 of 38 tracked vendors sell coding environments. That divergence shows which domains lack vendors.
  • Reading labor data at the task level has a research lineage, from OpenAI’s O*NET exposure scoring to Anthropic mapping millions of conversations onto occupational tasks.

Which marketplaces are tracked, and how are postings classified?

The three platforms serve different segments of the expert-supply market:

  • Mercor operates the largest expert network; its volume and margins are covered separately.
  • AfterQuery, from Y Combinator’s Winter 2025 batch, runs an expert platform supplying frontier labs with rubrics, environments, and computer-use trajectories.
  • micro1 focuses on expert roles for AI training and evaluation.

Every active posting is classified by domain, using each platform’s native taxonomy first and falling back to title-based inference only when the platform provides no domain tag. Each posting carries its source platform, native domain label, and the rule that assigned it, so misclassifications are auditable. Postings that match no domain rule are listed as unclassified with their title and platform, so the noise is visible.

Where does recruiting volume diverge from vendor supply?

In our reading of the postings active in June 2026 (counts not yet published), labs are recruiting experts across a broader set of domains than the vendor supply covers, particularly medical, legal, and finance roles, while vendor supply is concentrated in coding.

Stanford’s WORKBank survey shows the same mismatch from the employee side: it crossed 1,500 workers’ automation preferences with expert capability ratings across 844 tasks and 104 occupations, and found demand and capability rarely match (Shao et al., 2025). When a lab recruits a medical billing specialist through a marketplace, it is either building that environment internally or contracting a vendor to build it. Either way, the recruiting precedes the environment by months. Which of those jobs can become environments at all is a separate screen.

How reliable is marketplace data as a leading indicator?

OpenAI’s GPTs are GPTs scored ONET tasks (from the US Department of Labor’s occupational task database) for LLM exposure and set the precedent for treating the task, rather than the job title, as the unit of analysis. Anthropic’s Economic Index runs the same mapping in reverse, projecting millions of real Claude conversations onto ONET tasks to see where AI is already doing the work (Handa et al., 2025). Labor economists now extract O*NET structure from job-posting corpora at the scale of 155M listings (Meisenbacher et al., 2025). This research applies the same task-level lens to a narrower question: which expertise AI labs are paying to acquire in June 2026.

A lab posting for a payroll administrator or a dental biller is building an environment that requires domain expertise the lab does not have internally. The posting appears months before the environment reaches a training run, because environment construction, verifier calibration (checking that the program grading each episode agrees with experts), and expert validation all precede model exposure. The lag is our estimate from posting dates against launch dates, not a measured series.

What this means

Marketplace listings are the earliest public signal of where environment spend is heading. The domains where recruiting volume exceeds vendor supply are the domains where new environments will appear by the end of 2026, in our judgment.

FAQ

Why three marketplaces and not more?

Mercor, AfterQuery, and micro1 are the three marketplaces in our tracking set because each publishes expert postings with domain tags that isolate the AI-lab demand signal. Other lab-facing platforms were not included.

What is the classification noise problem?

Platform taxonomies are inconsistent: a role tagged “data” on one platform may be a finance role on another. Native taxonomy is used first and title-based inference as a fallback, with every posting’s classification rule recorded so misclassifications are auditable rather than invisible.

How far ahead of environment releases do postings appear?

Postings appear months before the environment reaches a training run. An expert recruited to author cases and calibrate a verifier is engaged before the environment exists. By the time a benchmark launches publicly, the recruiting that built it is months old; that lag is our estimate, not a measured series.