RL environments and AI benchmarks, cataloged daily.
The RL training market as universes you can fly through: every vendor, the RL environment builders, and the benchmarks they built. Each sits in the domain it leads with, and every placement links to the line on its own site or paper that put it there.
Snapshot 2026-10-10, rebuilt daily from public sources. How the data is built
How to read it
Star, based here: a vendor, in the realm of the domain it leads with. The logo is the vendor's; the glow is what it leads with selling.
Outpost: a ring on a realm's edge for a vendor based elsewhere that also covers that domain. A lane runs home.
Moons: the other things it sells. Hollow while it is still expanding into them.
Click any star for the lines from its own website that put it there.
Realms based here
The galaxy could not load. See every vendor on the board.
Today's signals · 2026-10-10
What changed in the catalog
October 10, 2026 was a score-wipe and quiet expansion day: 889 models had their the leaderboard overall scores cleared (set to empty), one new benchmark was added, seven existing leaderboards were updated, one new company joined the registry, and one new RL environment debuted. No models were added or removed, and tools, occupations, and research articles held flat.
Mass the leaderboard Score Wipe Across 889 Models
Every tracked model on the model registry had its overall score field cleared to an empty string today — from the top-ranked Claude Opus 5.5 (previously 86.32) and GPT-6 Astra (84.94) all the way down to GPT-5.2 (60.20). This is the largest single-day score wipe in the tracker's history and likely signals an imminent recalibration or methodology change on the leaderboard's end. Scores should be treated as unavailable until the next refresh.
One New Benchmark Lands
The benchmark count ticked up by one to 3,501. Details on the new entry were not surfaced in the representative diff, but it arrives alongside seven leaderboard row updates on existing benchmarks, continuing the steady week-over-week expansion of the benchmark registry.
New RL Environment Brings Total to 111
A single new environment was added today, pushing the environment count from 110 to 111. This continues a gradual but consistent growth trend: environments have risen from 98 in mid-August to 111 today, reflecting ongoing community contributions to the RL environment registry.
New Company Joins the Registry
One new company was added, bringing the total to 1,708. No metadata was surfaced in the diff, but the addition coincides with the new environment entry, suggesting the two may be related.
From the builders
What vendors are posting
The newest post on X from each RL environment vendor we track, refreshed with every build. Scroll the feed for more. How we find their accounts.
The catalog
One page per entry
Every environment, benchmark, domain and model has its own page with builder, domains and related work. Sort or filter any table. The snapshot rebuilds daily.
RL environments
Training and eval environments for agents, with builder and domains.
Browse → 80AI benchmarks
Public eval suites and leaderboards referenced across the ecosystem.
Browse → 24Domains
Coding, computer use, enterprise workflows, and more, each with its entries.
Browse → 0Model leaderboard
Ranked by overall leaderboard score, with creator and rank.
Browse →Domains by size
Ranked by total entries. 24 domains in all.
Show 18 smaller domains ↓ Hide smaller domains ↑
- 07 Math 4 env · 10 bench 14
- 08 Private Codebases 7 env · 2 bench 9
- 09 ML 7 env · 1 bench 8
- 10 Science 5 env · 2 bench 7
- 11 Finance 4 env · 0 bench 4
- 12 Creative 3 env · 0 bench 3
- 13 Knowledge 0 env · 3 bench 3
- 14 Security 3 env · 0 bench 3
- 15 Tool Use 3 env · 0 bench 3
- 16 Alignment 2 env · 0 bench 2
- 17 Browser 2 env · 0 bench 2
- 18 Chip Design 2 env · 0 bench 2
- 19 Healthcare 1 env · 1 bench 2
- 20 Audio 1 env · 0 bench 1
- 21 Games 1 env · 0 bench 1
- 22 Multilingual 0 env · 1 bench 1
- 23 Reasoning 0 env · 1 bench 1
- 24 Search 0 env · 1 bench 1