How the data is built
Every page is rebuilt each morning from public catalogs of RL environments and AI benchmarks, then cleaned, merged and categorized by our own synthesis pipeline.
Sources factored in
Four public sources are factored into the data rendered across the site. Their records are combined, deduplicated and shown in our own format, so individual pages do not cite them one by one.
- rl-list, a directory of companies building RL environments and the areas they focus on.
- Pavlov’s List, a tagged list of RL environments and benchmarks.
- BenchLM, a benchmark index and model leaderboard. Model scores on our leaderboard are its overall scores, carried over as published.
- BenchmarkList, a registry of AI benchmarks and the categories they cover.
We do not re-run any evaluation. A score on our site is the score the public leaderboard reported on the snapshot date.
What the pipeline does
- Scrape. A worker collects each source every morning and keeps the day’s snapshot, so any day can be compared with the one before it.
- Rank and deduplicate. Entries that appear in more than one source are deduplicated by name, keeping the strongest record. Placeholders are dropped, and each section keeps its most notable entries.
- Synthesize. Each source names things its own way (“Code”, “coding”, “Software engineering”). Our synthesis step maps every raw tag to one canonical category, and every raw location to a real city. An AI agent makes each new decision once, from sample entries, and the decision is cached in our repository so the vocabulary stays stable from day to day.
- Describe. Each category gets a one or two sentence definition. Editors wrote the first ones; a new category gets a draft definition from the same agent, grounded in the entries it holds.
- Brief. A daily briefing compares the snapshot with the previous day and summarizes what changed. It appears on the homepage under “What changed in the catalog”.
- Publish. The site, the slide and the JSON API are rebuilt from the result.
Vendor accounts on X
Each RL environment vendor is linked to its official account on X. We take the account the vendor links from its own website, or one an editor has checked when the site links none. Founders’ personal accounts are not used. Vendor pages and the homepage show each account’s newest posts, fetched with every build and shown with X’s official embed (with Do Not Track set, so X is asked not to personalise from it).
Logos and preview images are read from each vendor’s own website by the same daily run that collects the catalog, so a new vendor or benchmark gets its logo the day it appears. A benchmark’s come from its own site, its code repository or its publisher, never from a leaderboard that only reports its scores. Logos and images belong to their owners.
Snapshot dates
Every page shows the snapshot it was built from. The homepage checks for a newer snapshot when it loads and updates its counts and briefing if one is available.