Who pays for open AI agent benchmarks? A data vendor's $3M in-kind grants now back seven
By Bam AltmanParody pen name · View profile$3M in expert data, compute credits, and engineering across 7 benchmarks; 38.2% best pass rate on Agents' Last Exam's 147 public tasks

Open AI agent benchmarks are increasingly funded in kind by the vendors that sell evaluation data. Snorkel AI committed $3M to its Open Benchmarks Grants on February 11, 2026, delivered as expert data services, compute credits, and engineering time rather than cash, and the first cohort, announced in July 2026, backs seven benchmarks: Agents’ Last Exam, OSWorld 2.0, Continual Learning Bench, SlopCode Bench, Frontier-Bench, Terminal-Bench Science, and Senior SWE-Bench. The funded benchmarks are hard. On the 147 public tasks of Agents’ Last Exam hosted on Snorkel’s leaderboard, the best agent passes 38.2%, and the benchmark’s authors report an average full pass rate below 1% across the whole corpus. Alex Ratner, Snorkel’s co-founder and chief executive, presented the program in the Evals track at Runtime, Modal’s conference in San Francisco, on October 1, 2026. The arrangement places the supplier of expert data inside the benchmarks that measure what expert data is worth.
Key Takeaways
- The seven funded benchmarks and their homes: Agents’ Last Exam (UC Berkeley RDI), OSWorld 2.0 (HKU XLANG Lab), Continual Learning Bench and SlopCode Bench (University of Wisconsin-Madison), Frontier-Bench (Laude Institute and the Harbor community), Terminal-Bench Science, and Senior SWE-Bench (Princeton and Wisconsin-Madison).
- Grants are in kind, and outputs must be released under permissive licenses: MIT or Apache 2.0 for code, CC BY 4.0 or CC0 for datasets.
- Agents’ Last Exam holds 1,500+ expert-sourced tasks across 55 sub-industries from 300+ experts, toward a 5,000-task target; every frontier agent tested scored 0% on its hardest tier.
Which benchmarks did the program fund?
Snorkel announced the program on February 11, 2026 in a post by Vincent Sunn Chen, with Hugging Face, Prime Intellect, Together AI, Factory HQ, Harbor, and PyTorch as founding partners supplying research, advisory support, and platform credits. Applications opened March 1 on a rolling basis with quarterly selections, and the first cohort followed in July. The $3M is an estimated value of expert data-as-a-service access, compute credits, and Snorkel engineering support rather than a cash transfer. Recipients keep their intellectual property and must publish under permissive licenses.

Photo: rlresearch.ai, Runtime conference, The Midway, San Francisco, October 1, 2026
Ratner’s argument on stage was that the field has the problem backwards: “we actually have AI capabilities significantly leading our ability to evaluate them”, and what reads as saturation is a backlog: “it’s not that everything is saturated, it’s that we’re actually just behind on building these more sophisticated, more realistic, more precise evals, data sets and benchmarks”. His example was Senior SWE-Bench, built with the Princeton group behind SWE-bench, which moves toward “more complex context and also less precisely specified tasks”, still verifiable, with outputs “that are not just unit testable but that are also judged based on taste, architectural decisions, code quality, product decisions”.
How hard is Agents’ Last Exam?
Agents’ Last Exam (ALE), led by UC Berkeley’s RDI with Yiyou Sun as first author and more than 300 co-authors, was posted to arXiv on June 3, 2026. It organizes more than 1,000 tasks into 55 sub-fields and 13 industry clusters on the O*NET/SOC occupational taxonomy, sourced from 250+ industry experts, and reports an average full pass rate below 1% across the configurations tested. The RDI announcement puts the corpus at 1,500+ tasks from 300+ experts at 100+ institutions and says every frontier agent tested, Fable 5 included, scored 0% on the hardest tier.
Snorkel’s leaderboard scores a 147-task public subset with a 5,000-task target for the full corpus. Claude Opus 5.5 leads at a 38.2% pass rate, GPT-6 Astra follows at 34.2%, and GPT-6 Sol at 32.2%. ALE is built on the same occupational taxonomy as our coverage study, which found that 223 benchmarks definitively covered 3.5% of 202 O*NET tasks; a 5,000-task, expert-verified corpus on that taxonomy is the kind of benchmark that moves the number.
What does a vendor get from funding the ruler?
Snorkel sells expert data-as-a-service. The grants deliver that service to the labs building the benchmarks the industry grades against, and the license terms mean Snorkel owns none of the output. What it gains is position: the expert-sourced tasks pass through its data operation, its leaderboard hosts the public results, and its name sits on the page where frontier models are ranked. Two other vendors already own rulers. OpenAI’s GPT-5.6 release cited Surge’s GDP.pdf, and Mercor’s APEX-Agents patched and re-scored its own judge. Snorkel’s version differs in that the benchmarks are university-owned and permissively licensed, which makes them harder to bend and cheaper for competitors to use.
For environment vendors, a funded open benchmark sets the floor. A paid environment in a domain ALE covers has to measure something ALE’s free, expert-verified tasks cannot, or train a model past them. Ratner’s closing pitch to the room was that “there’s never been a more interesting time to actually build evals, build benchmarks”; for the vendor making it, every one of those benchmarks is a prospective buyer of expert data.
What this means
Benchmark funding has moved from labs and foundations to the vendors whose product the benchmarks consume, and the terms so far keep the outputs open. The test of the model is whether the funded benchmarks publish their verifier calibration as readily as their leaderboards.
FAQ
What is Agents’ Last Exam?
Agents’ Last Exam is an open benchmark from UC Berkeley’s RDI of long-horizon professional tasks with verifiable outcomes, sourced from 300+ industry experts across 55 sub-industries. The corpus exceeds 1,500 tasks toward a 5,000-task target; 147 are public and scored on Snorkel’s leaderboard, where the best pass rate is 38.2%.
Is the $3M paid in cash?
No. Snorkel describes the commitment as the estimated value of expert data-as-a-service access, compute credits, and dedicated engineering support. Recipients keep their intellectual property and must release code under MIT or Apache 2.0 and datasets under CC BY 4.0 or CC0.