Glossary
The words this site uses to sort the RL training market, and how a vendor ends up in each one.
What vendors sell
A vendor can sell several of these. Each one is either core (what the vendor leads with today) or expanding (newly announced, beta or a recent launch).
- RL environments 48 vendors →
- A simulated workplace where an AI agent practises a task (a code repository, a browser, business software) and an automatic grader scores each attempt. Labs train on thousands of these attempts. Bought to improve models. For example: Fleet, Plato, Mechanize, Deeptune.
- Evals 46 vendors →
- A fixed, held-out set of tasks with fixed scoring, used to measure and compare models. Never trained on, so the score means something. Used to measure models. For example: Vals AI, Patronus AI, Scale's leaderboards.
- Human data 25 vendors →
- A network of experts (engineers, doctors, lawyers, analysts) hired to write, judge or demonstrate tasks: the human judgement behind training data, graders and evals. For example: Mercor, Surge AI, micro1, Turing.
- Sandboxes and compute 8 vendors →
- The machines environments run on: isolated containers or virtual machines, started by the thousand, where an agent can safely run code or click through apps. They sell the runtime, not the tasks. For example: E2B, Daytona, Modal, Runloop.
Other offerings
- Datasets 26 vendors →
- Ready-made or licensed data sold as a product: post-training sets, agent trajectories, licensed corpora. Bought once, not commissioned task by task. For example: Wirestock, Exabite.
- Training platforms 4 vendors →
- Hosted services for running RL or post-training jobs and serving the resulting models. For example: Prime Intellect.
- Agent safety and verification 7 vendors →
- Red-teaming, guardrails and runtime checks that test or constrain what agents do once deployed. For example: Gray Swan, Align AI, Veris AI.
- Enterprise AI services 8 vendors →
- Engineers and applications deployed inside a company to build or run its AI systems. For example: Scale (applications), Mercor (forward-deployed engineers).
Domains
The subject matter of the tasks or data. A vendor lists every domain it covers and is shown under the one it leads with. A vendor that covers three or more with none leading is Multi-domain. Pure infrastructure has no domain.
- Coding
- Writing, reviewing and running software: repositories, terminals, tests.
- Computer use
- Operating a computer like a person: browsers, desktop apps, GUIs.
- Enterprise work
- Knowledge work in business software: CRM, ERP, email, spreadsheets, support, operations.
- Finance
- Banking, trading, accounting, insurance and financial analysis.
- Legal
- Legal research, contracts, compliance and regulation.
- Health and bio
- Medicine, clinical work, biology and drug discovery.
- Science
- Physical sciences and research workflows outside biology.
- Math
- Mathematical reasoning and proofs.
- Security
- Cybersecurity, capture-the-flag tasks and AI red-teaming.
- Games
- Video games and game-like simulations.
- Robotics
- Embodied and physical AI: robots, factories, manufacturing.
- Chip design
- Hardware and semiconductor design.
- Media
- Audio, voice, image, video and creative work.
- Multi
- Generalists: vendors covering three or more domains without one they lead with. Their specific domains are still listed.
Task traits
Properties of the tasks themselves, recorded when a vendor's site states them.
- Long-horizon
- Tasks that take many steps or hours to finish.
- Tool use
- Tasks that call tools or APIs, including MCP servers.
- Reasoning
- Tasks that test multi-step reasoning more than any one skill.
- Private data
- Built on private codebases or proprietary data.
- Multilingual
- Tasks in several human languages.
How vendors are placed
- We read each vendor's homepage and one product page from its own website.
- A model sorts the vendor into the fixed categories above, using these definitions. It can only pick from this list.
- Every offering needs a quote from the vendor's site. If the quote can't be found on the page, the offering is dropped.
- Pages are re-read monthly, sooner when they change. New offerings, dropped ones and moves into a new domain show up as dated changes on the landscape.
Categories version 2026-10-10.1.