Can AI invent its own training data? Adaption's zero-seed results move the scarce input to verification
By Spark BenioffParody pen name · View profile17% relative quality gain and 19% diversity gain across 8 task types, widening to 37% at 20,000 samples

A model can now write its own training set from a paragraph of description and beat five frontier APIs at the job, which moves the scarce input in post-training from examples to the graders that judge them. Adaption’s Invent-a-Dataset, a technical report posted to arXiv on October 1, 2026 by Shivalika Singh, Andrija Djurisic, Gbemileke Onilude, Sudip Roy, and Sara Hooker, starts from zero seed data and reports 17% higher quality and 19% higher diversity, relative, than datasets generated by Anthropic, Google, OpenAI, DeepSeek, and Z.ai models across eight task types. The diversity gap widens with scale, from parity at 200 samples to 37% at 20,000. Hooker, Adaption’s co-founder and chief executive, presented the work at Runtime, Modal’s conference in San Francisco, the same day. The result lowers the price of examples and leaves the price of verification where it was: deciding whether invented answers are correct still needs a verifier or a human.
Key Takeaways
- Invent-a-Dataset goes from a dataset description to a post-training dataset with no seed data, no schema, and no labeling guide; models fine-tuned on its output rank higher downstream across model architectures.
- AutoScientist, launched May 13, 2026, co-optimizes the data and the training recipe and reports a 35% average gain over human-configured training; Hooker said on stage that its per-vertical win rates had clustered near 60% because of how the objective was specified.
- Generation supplies inputs and reference answers. It does not supply the check on whether those answers are correct, and Adaption’s stated focus is non-verifiable problems, where that check is a person.
What did Adaption measure?
The report frames the zero data regime as the setting most practitioners face: a capability to teach and no examples of it. Invent-a-Dataset is a prompt-based system that expands a description into a large dataset, and the comparison runs it against five frontier model APIs asked to do the same across eight task types and dataset sizes up to 20,000 samples. Its samples score 17% higher on quality and 19% higher on diversity in relative terms, the diversity advantage grows from parity at 200 samples to 37% at 20,000, and models fine-tuned on its data rank consistently higher than models fine-tuned on the other generators’ output.

Photo: rlresearch.ai, Runtime conference, The Midway, San Francisco, October 1, 2026
Adaption’s product post of September 3, 2026 positions the system as the first half of a workflow: describe the behavior, generate the data, then hand both to AutoScientist to co-optimize data and training.
What is AutoScientist, and what did its author say about the chart?
AutoScientist, announced by Hooker on May 13, 2026, searches over training configurations and data mixtures against a stated objective and reports a 35% average gain over human-configured training. The slide at Runtime showed win rates against the original model across eight verticals, rising from a range of 31% to 41% before to 59% to 69% after.

Photo: rlresearch.ai, Runtime conference, The Midway, San Francisco, October 1, 2026
Hooker’s gloss added what the chart left out. The results clustered near 60% because of how the objective had been specified, she said, and after that specification changed the numbers rose: “your ability to improve is also so dependent on how you specify your objective”. She also located Adaption’s work precisely: “our focus is actually non-verifiable problems”, the ones where code and design made rapid progress only because they had tight feedback interfaces and good gold standards. Her prescription followed: “Instead of building the biggest model, you infer the type of problem and everything in your stack should change as quickly as possible.” And the consequence for model builders: “The best model now will be one that interacts.”
What does zero-seed generation not replace?
The generator writes inputs and reference outputs. The report scores those on quality and diversity, and the downstream check is whether a fine-tuned model ranks higher on held-out evaluations. None of that establishes that a given reference answer is correct, which is the property RL with verifiable rewards depends on. A reward computed from a wrong reference trains the wrong behavior with the same efficiency as a right one, and the reference-making model and the grading model often share a family, the self-referential loop our ground truth analysis warns about.
Surge’s 1,700-task study is the human-authored counterpart: expert-written tasks with checkable outcomes lifted five coding benchmarks at once. Invent-a-Dataset cuts the cost of the first column of that table, the task text. The second column, a verifier an expert signed off on, is unchanged in price, and it is the column environment vendors sell.
What this means
Examples are becoming an API-priced commodity, and the gains reported here say the commodity is good. The scarce input moves to whatever decides correctness: a deterministic verifier where one exists, an expert where it does not, and in Hooker’s own framing, an interface that gets a person’s judgment into the loop fast.
FAQ
What is a zero data regime?
The setting in which a team wants a model to learn a capability and has no examples of it, no seed corpus, no schema, and no labeling guide. Adaption’s report calls it the most extreme case practitioners face and the most common.
Does synthetic data make expert-built environments obsolete?
No. Generated data lowers the cost of task text and reference answers; an environment’s value is the verifier that checks an outcome and the expert adjudication behind it. The report measures generator quality and downstream rank, and leaves reference correctness unmeasured.
What is AutoScientist?
AutoScientist is Adaption’s system for co-optimizing a training recipe and its data against a stated objective. Announced May 13, 2026, it reports a 35% average gain over human-configured training; paired with Invent-a-Dataset it forms a path from objective to model with no starting data.