The aggregate view

Leaderboard

Environments, benchmarks, and model rankings. One place to explore the catalog.

Snapshot 2026-09-21 · Model scores retain their source rankings. Environments and benchmarks are catalogs, not scored rankings.

RL Environments Catalog

Dedicated view ↗
Domain
Source

80 environments shown

EnvironmentDomainsSource
AfterQueryAfterQuery is a San Francisco applied-research lab and data platform (YC W25) that supplies frontier AI labs…CodingComputer UseEnterprise Workflowsrl-list
MercorMercor is a venture-backed expert-marketplace and AI-training-data company that organizes a network of…CodingEnterprise WorkflowsLong-Horizonrl-list
Prime IntellectPrime Intellect operates an open-source RL stack - the Environments Hub (2,500+ community RL environments),…CodingEnterprise WorkflowsLong-HorizonMathrl-list
Surge AISurge AI is a bootstrapped, high-revenue human-data and RLHF labeling leader serving frontier AI labs, which…Enterprise WorkflowsLong-Horizonrl-list
MechanizeMechanize is a small, elite San Francisco vendor (founded April 2025 by ex-Epoch AI researchers Matthew…CodingPrivate Codebasesrl-list
ProximalProximal is a San Francisco-based (with a Bangalore presence) research lab for coding data, building…CodingLong-HorizonPrivate Codebasesrl-list
DeeptuneDeeptune was a New York-based startup building managed reinforcement-learning environments ('training gyms')…CodingComputer UseEnterprise WorkflowsPrivate Codebasesrl-list
Sepal AISepal AI was a YC-backed (S24) San Francisco data-research company that built high-quality training data,…Enterprise WorkflowsLong-HorizonMathrl-list
TuringTuring is a large AGI-infrastructure and engineering-services company that supplies frontier AI labs with…CodingEnterprise WorkflowsPrivate Codebasesrl-list
Fleet AIFleet AI builds high-fidelity reinforcement-learning training environments ('gyms') that replicate enterprise…Computer UseEnterprise Workflowsrl-list
Gray Swan AIGray Swan AI is a Pittsburgh-based AI security company spun out of Carnegie Mellon, offering adversarial…rl-list
RefreshRefresh (YC X25) builds simulation engines / RL environments with verifiable rewards for coding and computer…CodingComputer UsePrivate Codebasesrl-list
Scale AIScale AI is the data-labeling and AI-data incumbent that has extended into RL environments, offering simulated…CodingComputer UseEnterprise WorkflowsLong-Horizonrl-list
ModalModal (Modal Labs) is a New York-based, Python-native serverless cloud purpose-built for AI/ML workloads,…Codingrl-list
ScaleStanford, Coinbase, YCMulti-DomainCodingpavlovslist
SurgeGoogle, Meta, TwitterMulti-Domainpavlovslist
micro1Stanford, BerkeleyMulti-Domainpavlovslist
NormalCohere, Notion, FuseChip Designpavlovslist
FleetMercor, Anthropic, MSLMulti-DomainEnterprise Workflowspavlovslist
Bespoke LabsBespoke Labs is an applied AI research lab (Mountain View, CA, founded 2024) focused on data curation and…Long-Horizonrl-list
Vals AIVals AI is an independent, third-party benchmarking and evaluation platform that scores LLMs and AI…rl-list
MorphMorph (Morph Labs) provides snapshot-based VM compute for AI agents via its Infinibranch / Liquid Metal…Codingrl-list
Andon LabsAndon Labs is a Y Combinator-backed (W24) startup, formerly Vectorview, building benchmarks and evaluations…Computer UseLong-Horizonrl-list
HalluminateHalluminate (YC S25, founded 2024, San Francisco) builds managed reinforcement-learning sandbox environments,…Computer UseEnterprise Workflowsrl-list
BenchFlowBenchFlow is an early-stage, YC-backed open-source 'environment lab' building evaluation infrastructure and a…CodingComputer UseEnterprise Workflowsrl-list
General ReasoningGeneral Reasoning is an AI research lab (operating research hub in London; legal entity General Reasoning,…CodingLong-Horizonrl-list
Good Start LabsGood Start Labs is a 2025 Every spin-out that builds game-based environments to generate…Long-Horizonrl-list
Chakra LabsChakra Labs runs Dojo, an open/collaborative reinforcement-learning environment hub for computer-use agents,…Computer Userl-list
VmaxVmax is a San Francisco reinforcement-learning startup (founded 2025 by three RL/robotics PhDs from UCL and…CodingLong-Horizonrl-list
DatacurveDatacurve is a YC W24 commercial data vendor that supplies expert-curated frontier coding data, RLHF traces,…CodingPrivate Codebasesrl-list
HUDHUD (YC W25, formerly hud.so) is a platform for building reinforcement-learning environments and evaluations…Computer UseEnterprise Workflowsrl-list
CuaCua (trycua, YC X25) is open-source MIT-licensed infrastructure for computer-use agents, providing cloud and…Computer Userl-list
Huzzle LabsHuzzle Labs is the AI division of London-based talent platform Huzzle (founded ~2020 by Ingmar Klein, Parham…CodingComputer UseEnterprise WorkflowsLong-HorizonPrivate Codebasesrl-list
CollinearCollinear AI operates a 'Simulation Lab' (SimLab) that builds sandboxed, stateful RL environments simulating…CodingComputer UseEnterprise WorkflowsLong-Horizonrl-list
Veris AIVeris AI sells a high-fidelity simulation platform plus a production runtime that let enterprises train,…Enterprise Workflowsrl-list
AndromedeAndromede is an early-stage RL data lab that programmatically generates RL environments, tasks, and verifiers…Long-Horizonrl-list
PlatoPlato (plato.so, Plato Technologies, Inc.) builds simulated worlds for training and evaluating browser and…Computer UseEnterprise Workflowsrl-list
RunloopRunloop sells cloud-hosted, isolated micro-VM 'devboxes' plus blueprints, snapshots and benchmark/eval tooling…Codingrl-list
MatricesMatrices builds reinforcement-learning training environments for frontier AI labs to train agents that use…Computer Userl-list
AIChampAIChamp builds custom reinforcement-learning environments and 'Virtual Gym' simulations for training and…Enterprise WorkflowsLong-Horizonrl-list
Habitat IncHabitat Inc is an early-stage commercial vendor (2-10 employees, New York HQ) building reinforcement-learning…CodingComputer UseEnterprise Workflowsrl-list
DaytonaDaytona provides secure, elastic, programmatic sandboxes ('computers') that AI agents and developers can spin…CodingComputer Userl-list
E2BE2B provides open-source, secure cloud sandboxes (built on Firecracker microVMs) for running AI-generated code…CodingComputer Userl-list
HandshakePalantir, Meta, Scale AIMulti-Domainpavlovslist
SnorkelStanford, UW, MetaMulti-DomainCodingpavlovslist
LatchBerkeley, Google, AsimovSciencepavlovslist
Patronus AIMeta AI, Google, AmazonMulti-DomainCodingpavlovslist
ParetoStanford, Scale, TuringMulti-DomainTool Usepavlovslist
ApturaLazard, McKinsey, EFFinanceEnterprise Workflowspavlovslist
QuesmaElastic, Sumo Logic, DynatraceSecuritypavlovslist
ReasonCoreTuring, Aera Technology, OracleScienceCodingpavlovslist
AbundantWaymo, Google, MercorCodingpavlovslist
AviroYC, Princeton, Georgia TechCodingMLpavlovslist
EdotEnvG-Research, Etched, ETH ZurichLong-HorizonMLpavlovslist
EmulatedAWS, DeepMind, Prime IntellectCodingMLpavlovslist
MarkovYC, IIT MadrasComputer UseTool Usepavlovslist
MetaphiWaymo, Nuro, NetflixCodingEnterprise Workflowspavlovslist
pre.devPenn M&T, WhartonCodingLong-Horizonpavlovslist
Tacit LabsMosaicML, MSR, EFScienceLong-Horizonpavlovslist
UlamOxfordMathpavlovslist
OriginatorGoogle, Dataform, Pareto HoldingsComputer UseCodingpavlovslist
ParsewaveMeta, Red HatLong-Horizonpavlovslist
WirestockUtah State, MITCreativepavlovslist
Taste LabsExa, Palantir, Mercado LibreCreativepavlovslist
Verita AIHRT, Jane Street, MercorCreativepavlovslist
ARIMLABSAkamaiSecurityLong-Horizonpavlovslist
dmodelOpenAI, Google Brain, EleutherAIMLAlignmentpavlovslist
IdlerMeta, Microsoft, CornellCodingpavlovslist
IncalmoCMU, MSR, RSAC LabsSecuritypavlovslist
Preference ModelAnthropic, Stripe, DatologyMLCodingpavlovslist
SymbalMeta AI, General Catalyst, DeeptuneCodingEnterprise Workflowspavlovslist
AkharaMeta, MIT, CornellEnterprise WorkflowsCodingpavlovslist
AnthromindGoogle, Forward HealthHealthcareLong-Horizonpavlovslist
Diffuse Labsa16z, CMUMLLong-Horizonpavlovslist
DisseiWarwick, Goldman Sachs, HBSFinancepavlovslist
ExabiteStanford, MIT, MagicCodingpavlovslist
HillclimbDeepMind, BaseMathpavlovslist
Ooak DataSamsung, GoPuff, YCEnterprise Workflowspavlovslist
PhinityNVIDIA, AWS, StanfordChip Designpavlovslist
Rise Data LabsWhartonEnterprise Workflowspavlovslist

AI Benchmarks Catalog

Dedicated view ↗
Domain
Source

80 benchmarks shown

BenchmarkDomainsSource
SWE-bench ProMulti-DomainCodingpavlovslist
Terminal-Bench 2.0State-of-the-art set of difficult terminal-based tasksCodingbenchmarklist
Terminal-Bench 2.1State-of-the-art set of difficult terminal-based tasksCodingbenchmarklist
WebArenaBrowserGym leaderboard slice for WebArena, evaluating autonomous web agents across realistic browser tasks.Agenticbenchmarklist
OSWorldA computer-use benchmark for GUI task completion across the broader OSWorld task suite.Agenticbenchlm-benchmarks
SWE-benchVals AI’s independent run of the public SWE-bench issue set, reported by human time-to-fix bucket. Vals…Codingbenchlm-benchmarks
Terminal-Bench / Terminal-Bench 2.0Terminal-Bench / Terminal-Bench 2.0 (arXiv:2601.11868)CodingComputer UseEnterprise Workflowsrl-list
Senior SWE-BenchMulti-DomainCodingpavlovslist
OSWorld-VerifiedComputer UseEnterprise Workflowsrl-list
FinanceQAFinanceQA: A Benchmark for Evaluating Financial Analysis Capabilities in LLMs (arXiv:2501.18062;…CodingComputer UseEnterprise Workflowsrl-list
IDE-BenchCodingComputer UseEnterprise Workflowsrl-list
EnterpriseBench CoreCraftEnterpriseBench CoreCraft: Training Generalizable Agents on High-Fidelity RL Environments - arXiv:2602.16179…Enterprise WorkflowsLong-Horizonrl-list
APEXCodingEnterprise WorkflowsLong-Horizonrl-list
APEX-AgentsCodingEnterprise WorkflowsLong-Horizonrl-list
APEX-SWECodingEnterprise WorkflowsLong-Horizonrl-list
AIME 2025All 30 problems from the 2025 American Invitational Mathematics Examination, testing olympiad-level…Mathbenchmarklist
BrowseCompBrowseComp: Evaluates autonomous agent performance on multi-step tasks requiring planning, state tracking,…Agenticbenchmarklist
FHIR-AgentBenchFHIR-AgentBench evaluates LLM agents on realistic interoperable EHR question answering over HL7 FHIR…Healthcarebenchmarklist
FrontierMath 2025-02-28 PrivatePrivate FrontierMath research-level mathematics benchmark snapshot.Mathbenchmarklist
FrontierMath Tier 4 2025-07-01 PrivatePrivate Tier 4 FrontierMath problems at research-level mathematical difficulty.Mathbenchmarklist
Global MMLUGlobal MMLU multilingual knowledge-and-reasoning evaluation reported in Anthropic's Claude Opus 4.8 system…Multilingualbenchmarklist
GPQA DiamondThe hardest GPQA subset of graduate-level science questions in biology, chemistry, and physics.Reasoningbenchmarklist
GSM8KGrade-school math word-problem benchmark for evaluating multi-step arithmetic and reasoning performance.Mathbenchmarklist
Humanity's Last ExamFrontier-level benchmark with expert-vetted closed-ended questions across mathematics, sciences, and…Knowledgebenchmarklist
LiveCodeBench-PlusThe rapid progress of frontier large language models has led to widespread benchmark saturation, limiting the…Codingbenchmarklist
Long-Horizon Terminal-BenchLong-Horizon Terminal-Bench evaluates whether agents can sustain progress on 46 reproducible terminal tasks…Agenticbenchmarklist
MMLU-ProEnhanced MMLU benchmark with graduate-level questions across 14 subject areas and ten answer options.Knowledgebenchmarklist
MultiMedia-TerminalBenchTerminals provide a powerful interface for AI agents by exposing diverse tools for automating complex…Agenticbenchmarklist
OSWorld 2.0Anthropic first-attempt evaluation on the OSWorld authors August 2026 task release, preserving both partial…Computer Usebenchmarklist
OTIS Mock AIME 2024-2025Competition-level math problems from OTIS Mock AIME evaluating olympiad-level math.Mathbenchmarklist
SimpleQA Net ScoreSimpleQA closed-book factuality results reported by Anthropic as net score: correct responses minus incorrect…Knowledgebenchmarklist
SWE-bench DockerSWE-bench Docker: Evaluates software-engineering agents on realistic issue resolution, repository navigation,…Codingbenchmarklist
SWE-bench ExtraSWE-bench Extra: Evaluates software-engineering agents on realistic issue resolution, repository navigation,…Codingbenchmarklist
SWE-bench JavaScriptSWE-bench JavaScript: Evaluates software-engineering agents on realistic issue resolution, repository…Codingbenchmarklist
SWE-bench VerifiedBash-only variant of SWE-bench Verified for real-world GitHub issue resolution.Codingbenchmarklist
Terminal BenchTerminal and command-line interaction tasks for evaluating agent performance.Agenticbenchmarklist
Terminal-Bench HardArtificial Analysis Terminal-Bench hard subset for terminal-based software engineering, system administration,…Agenticbenchmarklist
AIMEChallenging national math exam given to top high-school studentsMathbenchmarklist
LiveCodeBenchOur Implementation of the LiveCodeBench benchmarkCodingbenchmarklist
Hy-BrowseComp-Pro2Tencent internal agentic-search evaluation reported in the HY4 preview benchmark appendix; distinct from…Searchbenchmarklist
SWE-bench Verified MiniHAL's cost-aware agent leaderboard for the SWE-bench Verified Mini software engineering subset.Codingbenchmarklist
TAU-bench AirlineHAL's standardized, cost-aware agent leaderboard for TAU-bench Airline customer-service tasks.Agenticbenchmarklist
HELM GSM8KHELM GSM8K: Measures mathematical reasoning, symbolic problem solving, proof construction, or…Mathbenchmarklist
DeepSWECodingpavlovslist
CoreCraftMulti-Domainpavlovslist
FinanceBenchMulti-DomainCodingpavlovslist
SWE-MarathonCodingpavlovslist
SWECodingMLpavlovslist
Terminal-Bench 3.0A continuously maintained benchmark for difficult computer work, including coding, deep learning, finance,…Agenticbenchlm-benchmarks
Terminal-Bench 4.0The current Terminal-Bench release measures difficult computer work after recalibrating task resources, fixing…Agenticbenchlm-benchmarks
Terminal-Bench-Science 0.1A benchmark of AI agents completing expert-curated research workflows across the life, physical, Earth,…Agenticbenchlm-benchmarks
HLE w/ toolsTool-augmented Humanity's Last Exam scores reported in DeepSeek-V4 thinking-mode evaluations.Agenticbenchlm-benchmarks
AA Terminal-Bench 4.0An independently evaluated Terminal-Bench v4.0 result from Artificial Analysis, one of the ten components of…Agenticbenchlm-benchmarks
BrowseComp-VLA vision-language browsing benchmark for multimodal web research and tool-use workflows.Agenticbenchlm-benchmarks
TAU-benchOriginal TAU-bench evaluates a model-driven agent in simulated airline and retail customer-service…Agenticbenchlm-benchmarks
WebArena-VerifiedWebArena-Verified is an audited release of the WebArena browser-agent benchmark. It rechecks task…Agenticbenchlm-benchmarks
AA LiveCodeBenchAn independently evaluated LiveCodeBench result from Artificial Analysis.Codingbenchlm-benchmarks
AA Terminal-Bench 2.1An independently evaluated Terminal-Bench v2.1 result from Artificial Analysis.Codingbenchlm-benchmarks
HumanEvalA set of 164 handwritten Python function-generation problems. HumanEval is useful as a historical floor check,…Codingbenchlm-benchmarks
CodeforcesCompetitive-programming rating reported for DeepSeek-V4 thinking-mode evaluations.Codingbenchlm-benchmarks
LiveCodeBench v6LiveCodeBench v6 is a named release slice used in provider comparison tables. Keeping it separate prevents v6…Codingbenchlm-benchmarks
LiveCodeBench v5LiveCodeBench v5 is a named release and date-window slice. BenchLM keeps explicitly labeled v5 rows outside…Codingbenchlm-benchmarks
LiveCodeBench Pass@1-COTThis lane contains DeepSeek's LiveCodeBench Pass@1-COT results. The explicit metric and prompting label keeps…Codingbenchlm-benchmarks
LiveCodeBench ProA harder competitive-programming benchmark family built from Codeforces, ICPC, and IOI problems, with…Codingbenchlm-benchmarks
Multi-SWE BenchA multi-language software-engineering benchmark that measures repository-level bug fixing and implementation…Codingbenchlm-benchmarks
ARC-AGI-1ARC Prize fluid-intelligence benchmark using novel visual grid transformations.Reasoningbenchlm-benchmarks
ARC-AGI-2A benchmark measuring fluid intelligence and novel abstract reasoning through visual grid puzzles. Models must…Reasoningbenchlm-benchmarks
ARC-AGI-3An interactive successor to ARC-AGI-2 that evaluates whether an AI agent can learn unfamiliar task mechanics…Reasoningbenchlm-benchmarks
SWE-bench MultimodalA multimodal variant of SWE-bench that adds visual context (screenshots, design mockups) to software…Multimodalbenchlm-benchmarks
AA MMLU-ProAn independently evaluated MMLU-Pro result from Artificial Analysis.Knowledgebenchlm-benchmarks
MMLUA comprehensive multiple-choice question answering test covering 57 tasks including elementary mathematics, US…Knowledgebenchlm-benchmarks
GPQAA challenging dataset of 448 multiple-choice questions written by domain experts in biology, physics, and…Knowledgebenchlm-benchmarks
GPQA-DA display-only GPQA Diamond reference from provider comparison charts.Knowledgebenchlm-benchmarks
HLEAn expert-authored benchmark designed to probe frontier knowledge and reasoning. BenchLM keeps protocol…Knowledgebenchlm-benchmarks
HLE-VerifiedA verified and revised Humanity's Last Exam set that removes uncertain items and repairs fixable questions…Knowledgebenchlm-benchmarks
AA-GPQA DiamondA display-only Artificial Analysis GPQA Diamond score.Knowledgebenchlm-benchmarks
AA-HLEA display-only Artificial Analysis Humanity's Last Exam score.Knowledgebenchlm-benchmarks
SimpleQAA benchmark that evaluates the ability of language models to answer short, fact-seeking questions accurately.…Knowledgebenchlm-benchmarks
Chinese-SimpleQAA Chinese short-form factuality benchmark reported by DeepSeek for V4 model evaluations.Knowledgebenchlm-benchmarks
HLE w/o toolsTool-free variant of Humanity's Last Exam that isolates a model's raw frontier reasoning.Knowledgebenchlm-benchmarks

Model rankings

Dedicated view ↗

50 models shown

RankModelScoreCreatorSource
1Claude Fable 5.184.74Anthropicbenchlm-models
2GPT-6 Astra82.81OpenAIbenchlm-models
3Claude Opus 581.87Anthropicbenchlm-models
4Claude Fable 581.41Anthropicbenchlm-models
5GPT-5.6 Sol80.46OpenAIbenchlm-models
6Gemini 3.8 Flash75.89Googlebenchlm-models
7Kimi K374.43Moonshot AIbenchlm-models
8Qwen3.8 Max73.17Alibababenchlm-models
9Claude Opus 4.872.17Anthropicbenchlm-models
10GPT-5.571.86OpenAIbenchlm-models
11GPT-5.470.88OpenAIbenchlm-models
12Claude Opus 4.770.60Anthropicbenchlm-models
13Gemini 3.1 Pro70.14Googlebenchlm-models
14Grok 4.669.74xAIbenchlm-models
15Claude Sonnet 569.68Anthropicbenchlm-models
16Claude Opus 4.669.21Anthropicbenchlm-models
17Gemini 3.7 Flash68.52Googlebenchlm-models
18Gemini 3.6 Flash68.42Googlebenchlm-models
19Gemini 3.5 Flash68.02Googlebenchlm-models
20Grok 4.567.99xAIbenchlm-models
21Grok 4.2067.27xAIbenchlm-models
22GLM-5.366.93Z.AIbenchlm-models
23Qwen3.7 Max66.87Alibababenchlm-models
24GLM-5.266.68Z.AIbenchlm-models
25Gemini 3 Pro66.36Googlebenchlm-models
26Kimi K2.7 Code65.60Moonshot AIbenchlm-models
27Kimi K2.665.50Moonshot AIbenchlm-models
28Ornith-1.5-397B65.32Ornith AIbenchlm-models
29GPT-5.264.86OpenAIbenchlm-models
30Qwen3.8-27B64.29Alibababenchlm-models
31dots3-note Preview64.14Dots Studiobenchlm-models
32GPT-5.3 Codex64.03OpenAIbenchlm-models
33Qwen 3.6 Max63.94Alibababenchlm-models
34DeepSeek V4 Pro 081363.69DeepSeekbenchlm-models
35GLM-5.163.26Z.AIbenchlm-models
36GPT-5.163.18OpenAIbenchlm-models
37Claude Sonnet 4.662.86Anthropicbenchlm-models
38MiMo-V2-Pro62.51Xiaomibenchlm-models
39Gemini 3 Flash61.94Googlebenchlm-models
40Qwen3.7 Plus61.60Alibababenchlm-models
41GLM-561.47Z.AIbenchlm-models
42MiniMax M361.25MiniMaxbenchlm-models
43Hy4 preview60.89Tencentbenchlm-models
44Hy360.79Tencentbenchlm-models
45Inkling60.38Thinking Machines Labbenchlm-models
46Qwen3.6 Plus60.36Alibababenchlm-models
47Grok 4.359.73xAIbenchlm-models
48Gemini 3.5 Flash-Lite58.78Googlebenchlm-models
49Claude Opus 4.558.42Anthropicbenchlm-models
50Grok 4.157.99xAIbenchlm-models