market Updated 4 min read

The RL environment industry is 38 companies, mostly under 50 people

31 of 38 tracked vendors have 50 or fewer employees

31 of 38 RL environment vendors have 50 or fewer employees; only Scale AI and Turing exceed 200

Our census of 38 tracked RL environment vendors finds 31 with 50 or fewer employees. The companies supplying the most capitalized labs on earth are a cottage industry: Scale AI and Turing are the only firms above 200 staff, and everyone else, including the vendors whose benchmarks appear in frontier system cards (the documents labs publish with each model release), runs on a small team. On June 8, 2026, AfterQuery, a 51-200-person company from Y Combinator’s Winter 2025 batch, published a +21.4% net win-loss gain (wins minus losses as a share of pairwise comparisons) on GDPval, OpenAI’s benchmark of professional work tasks, using on-policy distillation, a method that trains a student model on its own outputs with a teacher model’s feedback. The modal band is 11 to 50 employees, and both of the two largest vendors are legacy data-labeling companies that expanded into environments rather than companies founded to build them.

Key Takeaways

  • 31 of 38 tracked RL environment vendors have 50 or fewer employees; only Scale AI and Turing top 200.
  • AfterQuery moved GDPval by +21.4% net win-loss with on-policy distillation on June 8, 2026: a Qwen3.5-9B student distilled from a 397B-parameter teacher on AfterQuery’s task set.
  • The concentration risk affects vendors and labs alike: one lost lab contract can erase a vendor, and one acquisition can reset the market.

What does the headcount distribution look like?

Headcount distribution of 38 tracked RL environment vendors

The modal band is 11 to 50 employees: large enough to field a few domain teams, build a benchmark, and serve one to three lab customers, and small enough that a single acquisition resets the competitive landscape.

The two firms above 200 are Scale AI and Turing.

Who are the sub-50-person vendors?

Three vendors illustrate the pattern:

AfterQuery sits at 51-200, just above the sub-50 band. The company publishes FinanceQA and IDE-Bench, supplies expert-generated rubrics and environments to frontier labs, and reported a 5x uplift on Terminal-Bench 2.0 (3.1% to 17.0%) from its expert data. Its size is unusual among tracked vendors.

How do small vendors prove value?

Small vendors earn lab contracts with a measured uplift on a benchmark the lab already tracks. AfterQuery’s June 8 result is the example: a +21.4% net win-loss gain on GDPval from on-policy distillation. The method itself is public research, the generalized knowledge distillation (GKD) recipe from Agarwal et al. (2023). What the vendor adds is the task set and teacher setup the method runs on.

The academic record supports this. LIMO elicited sophisticated reasoning from a few hundred curated examples, and Stanford’s s1 matched much larger training runs with a curated set of 1,000. Both results show curated data beating volume, which a 20-person expert shop can supply.

Across the census, vendors winning lab spend point to a measured result: a benchmark number, a training uplift, a system-card citation. Headcount does not enter into it, and neither does revenue; the margins on this work are those of a services business at any size.

What this means

31 of 38 vendors are small enough that a single lost lab contract or a single acquisition removes them from the market. Acquirers find talent and tooling cheapest in the sub-50 band; a lab depending on one of those vendors is one failed quarter from losing it.

FAQ

Why are most RL environment vendors so small?

By our count the market has four to six buyers: the frontier labs. Those labs also build environments in-house, so vendors supplement internal teams rather than replace them. A vendor needs enough staff to build environments and grade them, but the demand does not support scaling beyond a handful of lab contracts.

What is GDPval?

GDPval is OpenAI’s benchmark of expert-graded professional work tasks across 44 occupations in the top GDP sectors (Patwardhan et al., 2025). Vendors publish training and evaluation results on it to show that their environments or data improve frontier-model performance on occupational work.

Does the headcount distribution predict acquisition targets?

Consolidation in this market will start with sub-50-person vendors, in our judgment: the 200+ firms are too expensive for most buyers and too entrenched to sell, and the sub-50 band is where talent and tooling cost least.