market Updated 5 min read

Why RL environment vendors earn services margins and software valuations

27% gross margin on Mercor's ~$2B in gross annualized revenue

Mercor keeps a 27% gross margin on roughly $2B flowing from AI labs to experts; the margin sits at the verification step

Frontier labs are shifting spend from labeled data to executable environments, and the vendor margins show where the leverage sits. Mercor, the expert marketplace that pays contracted professionals to build and grade training data for AI labs, runs roughly $2B in gross annualized revenue at a 27% gross margin, and is valued at $10B after a $350M Series C. AfterQuery closed a $30M Series A at a $300M valuation and, in the same April 9, 2026 announcement, reported passing a $100M revenue run rate. Mercor’s margin profile looks nothing like software: software gross margins typically run 70 to 80%, and a 27% margin is the cost of paying an expert for each unit of graded work. The category winner sells expert labor organized around verification, and the labs buying it cite vendor evals in their own system cards, the documents labs publish with each model release (June 2026). The money is moving from annotation budgets to environment budgets, toward the companies that can grade what a model does.

Key Takeaways

  • Mercor processes about $2B in gross annualized revenue at a 27% gross margin (its 2025 figure), and the market prices that services business like infrastructure: $10B post-money.
  • AfterQuery raised a $30M Series A at a $300M valuation and reported a $100M revenue run rate in the same April 2026 announcement.
  • On June 9, 2026 Anthropic cited Surge AI’s GDP.pdf and Riemann-bench in the Fable 5 system card; on June 2 Microsoft used Surge human evals for MAI-Thinking-1.

Update, September 16, 2026. The Mercor volume and margin figures were added after this post’s publication date. The Information reported on July 6, 2026 that Mercor crossed $2B in gross annualized revenue in June 2026, and on July 21 that its 2025 gross margin was 27%, rising to 33% in the second quarter of 2026 (as relayed by BigGo Finance).

What do the margins say about the business?

Mercor’s 27% gross margin is a services margin. The company processes roughly $2B in gross annualized revenue, money flowing from labs to contracted experts, and keeps about 27 cents of each gross dollar after paying the experts. A $10B valuation on that base prices in one belief: environment demand grows with every model release. In our reading it does not price in a cost structure that scales sublinearly, because Mercor has not shown one. The research on annotation labor says quality tracks pay and instruction clarity (Laux et al., 2023); headcount was not a variable in that study.

AfterQuery’s shape matches: a $30M Series A at a $300M valuation, a self-reported $100M run rate announced alongside it, and a product line of expert-generated rubrics, agent environments, and computer-use trajectories sold to frontier labs. Both companies are priced as infrastructure for a recurring demand cycle rather than as one-off data vendors.

What did Mercor’s thesis essay say?

Mercor published “The Economy will Become an RL Environment Machine” in September 2025, arguing that the market for humans teaching models is sized by what models cannot yet do. As capability advances, the work shifts from labeling correct answers to building the environments in which correct behavior is defined. The essay’s central claim: the addressable market is the gap between model competence and occupational competence, and that gap is where verification becomes the product.

RLVR, reinforcement learning with verifiable rewards, swaps learned reward models for deterministic checkers (Lambert et al., 2024), and DeepSeek-R1 showed frontier reasoning emerging from RL against verifiable signals rather than piles of human-labeled data. If the trainable signal is verification, the scarce input is whoever can define and grade correctness. The margin data is consistent with that: if the product were labeled data, margins would compress as labeling commoditizes. If the product is verification, margins hold as tasks get harder. Who defines that ground truth is the follow-on question.

Are vendors becoming frontier infrastructure?

On June 2 and June 9, 2026, two labs leaned on vendor evals in public:

When a lab cites a vendor’s eval in a system card, the vendor is supplying part of the frontier’s measurement layer. The spend that used to go to annotation vendors now goes to companies that can build and grade executable environments, and labs keep raising the stakes: Meta’s ScaleRL study used 400,000 GPU-hours to establish how RL post-training compute scales (Khatri et al., 2025). Each such scaling result raises the number of environments a lab needs to buy. The 27% margin shows the work is still labor-intensive; the system-card citations show labs now depend on the output. The vendors supplying it are mostly small companies.

Mercor's APEX-SWE chart: Claude Fable 5 at 65.5% Pass@1, ahead of Claude Opus 4.8 at 45.3%

A vendor benchmark used as the public scoreboard for a frontier release. This is the measurement layer the post argues labs are buying.

Mercor (@mercor) · June 9, 2026 · on X

What this means

Verification is where margin accrues. Labs buying environments are buying the ability to distinguish competence from performance, and the vendors that can do that at scale are the ones capturing the spend.

FAQ

Why is 27% gross margin a services margin?

Software businesses typically run 70 to 80% gross margins because serving one more customer costs almost nothing. Mercor’s 27% reflects the cost of paying experts for each unit of work. The marginal cost does not collapse, which defines a services business.

What is GDP.pdf?

GDP.pdf is a Surge AI benchmark for professional document tasks. Anthropic cited it in the Fable 5 system card (June 9, 2026) as an external measure of professional-work competence.

If margins are services-level, why are valuations software-level?

Valuations are software-level because every model release manufactures demand for the next batch of environments. Revenue recurs even though contracts do not: a new frontier model ships, the post-training position resets, and a fresh order follows.