market 5 min read

Why expert-data vendors are becoming research labs: 14 studies from four vendors in 40 days

14 vendor research posts in 40 days; Mercor Research announced Oct 8; Snorkel at a $375M run rate, up 18x in 12 months, valued at $3.5B

14 vendor research posts between September 1 and October 9, 2026, clustering toward the end of the window

The companies that sell expert data and reinforcement learning (RL) environments now publish research the way frontier labs do, and the cadence is the business model. Between September 1 and October 9, 2026, Mercor, Surge AI, Prime Intellect, and Snorkel put out 14 studies, benchmark launches, and methods write-ups, from a 397B training guide to a staff-engineer benchmark published two days ago. On October 8 Mercor announced Mercor Research, a team spanning evaluation, model training, economics, and labor, staffed by researchers from OpenAI, Sierra, Anthropic, Ai2, Bridgewater, and D.E. Shaw. Snorkel, which calls itself a frontier AI data lab, told TechCrunch on September 22 that its annualized run rate is $375M, up eighteenfold in a year, at a $3.5B valuation. Each study is a demonstration: the vendor’s tasks moved a model, or its benchmark caught one failing. The research is the sales motion, and the labs buying environments can read it as such.

Key Takeaways

  • The 14, by vendor: Surge 6 (Kimi hill-climbing, alignment baselines, DAYJOB, GDP.pdf post-training, GDP.xlsx, sudo L7), Mercor 4 (397B guide, APEX-Agents 1.1, 20x eval throughput, human baselines), Prime Intellect 3 (Goodfire case study, swarm essay, Rust rewrite), Snorkel 1 (environment design guide).
  • Mercor Research, announced October 8, runs from evals through post-training; its published result is Qwen3.5-397B’s Pass@1 on APEX-Agents rising from 16% to 27%. Its hires are covered separately.
  • Snorkel’s $350M Series E, led by Insight Partners and S32, values it at $3.5B, up from $1.3B seventeen months earlier; it books expert payments in cost of goods sold because it sells environments and datasets rather than labor.

What did the vendors publish in 40 days?

Fourteen vendor research posts between September 1 and October 9, 2026, by date and vendor

The count excludes product announcements and customer case studies and keeps anything with a method or a measurement. Mercor opened the window on September 1 with its 397B RL training guide and followed with APEX-Agents 1.1, an engineering note on a 20x gain in evaluation throughput, and the human baselines study. Surge published six: the 1,700-task Kimi study, its account of building human baselines for Anthropic’s alignment agents, DAYJOB, the GDP.pdf transfer study, GDP.xlsx, and sudo L7. Prime Intellect published the Goodfire case study on catching reward hacking with activation probes, the swarm essay, and the Rust rewrite. Snorkel published its environment design guide.

Six of the 14 report a trained model moving on a held-out benchmark, four launch or revise a benchmark with a leaderboard, and four describe methods. That mix is what a lab’s research blog looks like, with one difference: every result is also evidence for a product.

Why is Mercor building a research team now?

Mercor’s own explanation is a feedback loop between measurement and training. “APEX shows us where a model struggles.” “Experts help us understand why and produce training data around those weaknesses.” The team’s agenda covers evaluation, post-training, economics, and labor, and its first published proof is the guide that lifted Qwen3.5-397B from 16% to 27% Pass@1 on APEX-Agents with script, weights, and traces released. Who it hired, the first authors of LoRA, τ²-bench, SciCode, CritPt, and APEX-Agents and a co-lead of PostTrainBench, is the subject of our October 9 post. The timing follows the money: our economics analysis argued that vendor margins depend on verification rather than labor, and a research team is how a vendor proves its verification works, with a benchmark the labs cite and a training result they can reproduce. Mercor’s gross annualized revenue is $2B by TechCrunch’s reporting; the research team is the part of the company that turns that into evidence.

What does Snorkel’s balance sheet say about the model?

Snorkel’s September raise is the clearest statement of the economics. Its $375M run rate grew eighteenfold in twelve months, and the company told TechCrunch that because it sells RL environments and complete datasets rather than human labor, payments to its experts sit in cost of goods sold and are not reflected in that headline figure. TechCrunch’s own reporting puts the share such companies pay to domain specialists at roughly 60% to 70% of top-line income. A vendor that reports net revenue and books experts as COGS is describing itself as a product company, and a product company needs a research function to show the product works. Snorkel’s guide, its $3M in benchmark grants, and its hosted leaderboards are that function.

Prime Intellect’s version is infrastructure rather than data, and its demonstration is the same shape. The Rust rewrite ran 2,209 agents across 10,000 sandboxes with an objective verifier and no human grading, and its authors wrote the design rule a lab would: “We deliberately gave the loop no numeric targets, since a fixed threshold tends to become a stopping point.”

What this means

Four vendors have become publishers of first-party research, and a buyer should read each study as a product demonstration with the usual terms: the vendor chose the model, the tasks, and the judge. The studies that survive that reading, held-out benchmarks and released weights, are the ones that make a vendor worth a frontier contract.

FAQ

Which 14 posts are counted?

Mercor: 397B training guide (Sep 1), APEX-Agents 1.1 (Sep 8), 20x eval throughput (Sep 30), human baselines (Oct 1). Surge: Kimi K2.7 hill-climbing (Sep 11), Anthropic alignment baselines (Sep 15), DAYJOB (Sep 23), GDP.pdf post-training (Sep 29), GDP.xlsx (Sep 30), sudo L7 (Oct 8). Prime Intellect: Goodfire case study (Sep 17), On the Nature of the Swarm (Oct 6), Rust rewrite (Oct 9). Snorkel: environment design guide (Sep 28). Product launches and customer case studies are excluded.

What is Mercor Research?

A team Mercor announced on October 8, 2026 covering AI evaluation, model training, economics, and labor, staffed by researchers from OpenAI, Sierra, Anthropic, Ai2, Bridgewater, and D.E. Shaw, including LoRA first author Edward J. Hu. Its stated loop is that APEX benchmarks locate model weaknesses and experts produce training data against them.

Why does Snorkel book expert payments as cost of goods sold?

Because, by its account to TechCrunch, it sells RL environments and complete datasets rather than human labor, so the experts who build them are a production cost. The effect is that its $375M run rate is a net figure, unlike the gross figures some competitors report.