The aggregate view
Leaderboard
Environments, benchmarks, and model rankings. One place to explore the catalog.
Snapshot 2026-09-21 · Model scores retain their source rankings. Environments and benchmarks are catalogs, not scored rankings.
RL Environments Catalog
Dedicated view ↗| Environment | Domains | Source |
|---|---|---|
| AfterQueryAfterQuery is a San Francisco applied-research lab and data platform (YC W25) that supplies frontier AI labs… | CodingComputer UseEnterprise Workflows | |
| MercorMercor is a venture-backed expert-marketplace and AI-training-data company that organizes a network of… | CodingEnterprise WorkflowsLong-Horizon | |
| Prime IntellectPrime Intellect operates an open-source RL stack - the Environments Hub (2,500+ community RL environments),… | CodingEnterprise WorkflowsLong-HorizonMath | |
| Surge AISurge AI is a bootstrapped, high-revenue human-data and RLHF labeling leader serving frontier AI labs, which… | Enterprise WorkflowsLong-Horizon | |
| MechanizeMechanize is a small, elite San Francisco vendor (founded April 2025 by ex-Epoch AI researchers Matthew… | CodingPrivate Codebases | |
| ProximalProximal is a San Francisco-based (with a Bangalore presence) research lab for coding data, building… | CodingLong-HorizonPrivate Codebases | |
| DeeptuneDeeptune was a New York-based startup building managed reinforcement-learning environments ('training gyms')… | CodingComputer UseEnterprise WorkflowsPrivate Codebases | |
| Sepal AISepal AI was a YC-backed (S24) San Francisco data-research company that built high-quality training data,… | Enterprise WorkflowsLong-HorizonMath | |
| TuringTuring is a large AGI-infrastructure and engineering-services company that supplies frontier AI labs with… | CodingEnterprise WorkflowsPrivate Codebases | |
| Fleet AIFleet AI builds high-fidelity reinforcement-learning training environments ('gyms') that replicate enterprise… | Computer UseEnterprise Workflows | |
| Gray Swan AIGray Swan AI is a Pittsburgh-based AI security company spun out of Carnegie Mellon, offering adversarial… | ||
| RefreshRefresh (YC X25) builds simulation engines / RL environments with verifiable rewards for coding and computer… | CodingComputer UsePrivate Codebases | |
| Scale AIScale AI is the data-labeling and AI-data incumbent that has extended into RL environments, offering simulated… | CodingComputer UseEnterprise WorkflowsLong-Horizon | |
| ModalModal (Modal Labs) is a New York-based, Python-native serverless cloud purpose-built for AI/ML workloads,… | Coding | |
| ScaleStanford, Coinbase, YC | Multi-DomainCoding | |
| SurgeGoogle, Meta, Twitter | Multi-Domain | |
| micro1Stanford, Berkeley | Multi-Domain | |
| NormalCohere, Notion, Fuse | Chip Design | |
| FleetMercor, Anthropic, MSL | Multi-DomainEnterprise Workflows | |
| Bespoke LabsBespoke Labs is an applied AI research lab (Mountain View, CA, founded 2024) focused on data curation and… | Long-Horizon | |
| Vals AIVals AI is an independent, third-party benchmarking and evaluation platform that scores LLMs and AI… | ||
| MorphMorph (Morph Labs) provides snapshot-based VM compute for AI agents via its Infinibranch / Liquid Metal… | Coding | |
| Andon LabsAndon Labs is a Y Combinator-backed (W24) startup, formerly Vectorview, building benchmarks and evaluations… | Computer UseLong-Horizon | |
| HalluminateHalluminate (YC S25, founded 2024, San Francisco) builds managed reinforcement-learning sandbox environments,… | Computer UseEnterprise Workflows | |
| BenchFlowBenchFlow is an early-stage, YC-backed open-source 'environment lab' building evaluation infrastructure and a… | CodingComputer UseEnterprise Workflows | |
| General ReasoningGeneral Reasoning is an AI research lab (operating research hub in London; legal entity General Reasoning,… | CodingLong-Horizon | |
| Good Start LabsGood Start Labs is a 2025 Every spin-out that builds game-based environments to generate… | Long-Horizon | |
| Chakra LabsChakra Labs runs Dojo, an open/collaborative reinforcement-learning environment hub for computer-use agents,… | Computer Use | |
| VmaxVmax is a San Francisco reinforcement-learning startup (founded 2025 by three RL/robotics PhDs from UCL and… | CodingLong-Horizon | |
| DatacurveDatacurve is a YC W24 commercial data vendor that supplies expert-curated frontier coding data, RLHF traces,… | CodingPrivate Codebases | |
| HUDHUD (YC W25, formerly hud.so) is a platform for building reinforcement-learning environments and evaluations… | Computer UseEnterprise Workflows | |
| CuaCua (trycua, YC X25) is open-source MIT-licensed infrastructure for computer-use agents, providing cloud and… | Computer Use | |
| Huzzle LabsHuzzle Labs is the AI division of London-based talent platform Huzzle (founded ~2020 by Ingmar Klein, Parham… | CodingComputer UseEnterprise WorkflowsLong-HorizonPrivate Codebases | |
| CollinearCollinear AI operates a 'Simulation Lab' (SimLab) that builds sandboxed, stateful RL environments simulating… | CodingComputer UseEnterprise WorkflowsLong-Horizon | |
| Veris AIVeris AI sells a high-fidelity simulation platform plus a production runtime that let enterprises train,… | Enterprise Workflows | |
| AndromedeAndromede is an early-stage RL data lab that programmatically generates RL environments, tasks, and verifiers… | Long-Horizon | |
| PlatoPlato (plato.so, Plato Technologies, Inc.) builds simulated worlds for training and evaluating browser and… | Computer UseEnterprise Workflows | |
| RunloopRunloop sells cloud-hosted, isolated micro-VM 'devboxes' plus blueprints, snapshots and benchmark/eval tooling… | Coding | |
| MatricesMatrices builds reinforcement-learning training environments for frontier AI labs to train agents that use… | Computer Use | |
| AIChampAIChamp builds custom reinforcement-learning environments and 'Virtual Gym' simulations for training and… | Enterprise WorkflowsLong-Horizon | |
| Habitat IncHabitat Inc is an early-stage commercial vendor (2-10 employees, New York HQ) building reinforcement-learning… | CodingComputer UseEnterprise Workflows | |
| DaytonaDaytona provides secure, elastic, programmatic sandboxes ('computers') that AI agents and developers can spin… | CodingComputer Use | |
| E2BE2B provides open-source, secure cloud sandboxes (built on Firecracker microVMs) for running AI-generated code… | CodingComputer Use | |
| HandshakePalantir, Meta, Scale AI | Multi-Domain | |
| SnorkelStanford, UW, Meta | Multi-DomainCoding | |
| LatchBerkeley, Google, Asimov | Science | |
| Patronus AIMeta AI, Google, Amazon | Multi-DomainCoding | |
| ParetoStanford, Scale, Turing | Multi-DomainTool Use | |
| ApturaLazard, McKinsey, EF | FinanceEnterprise Workflows | |
| QuesmaElastic, Sumo Logic, Dynatrace | Security | |
| ReasonCoreTuring, Aera Technology, Oracle | ScienceCoding | |
| AbundantWaymo, Google, Mercor | Coding | |
| AviroYC, Princeton, Georgia Tech | CodingML | |
| EdotEnvG-Research, Etched, ETH Zurich | Long-HorizonML | |
| EmulatedAWS, DeepMind, Prime Intellect | CodingML | |
| MarkovYC, IIT Madras | Computer UseTool Use | |
| MetaphiWaymo, Nuro, Netflix | CodingEnterprise Workflows | |
| pre.devPenn M&T, Wharton | CodingLong-Horizon | |
| Tacit LabsMosaicML, MSR, EF | ScienceLong-Horizon | |
| UlamOxford | Math | |
| OriginatorGoogle, Dataform, Pareto Holdings | Computer UseCoding | |
| ParsewaveMeta, Red Hat | Long-Horizon | |
| WirestockUtah State, MIT | Creative | |
| Taste LabsExa, Palantir, Mercado Libre | Creative | |
| Verita AIHRT, Jane Street, Mercor | Creative | |
| ARIMLABSAkamai | SecurityLong-Horizon | |
| dmodelOpenAI, Google Brain, EleutherAI | MLAlignment | |
| IdlerMeta, Microsoft, Cornell | Coding | |
| IncalmoCMU, MSR, RSAC Labs | Security | |
| Preference ModelAnthropic, Stripe, Datology | MLCoding | |
| SymbalMeta AI, General Catalyst, Deeptune | CodingEnterprise Workflows | |
| AkharaMeta, MIT, Cornell | Enterprise WorkflowsCoding | |
| AnthromindGoogle, Forward Health | HealthcareLong-Horizon | |
| Diffuse Labsa16z, CMU | MLLong-Horizon | |
| DisseiWarwick, Goldman Sachs, HBS | Finance | |
| ExabiteStanford, MIT, Magic | Coding | |
| HillclimbDeepMind, Base | Math | |
| Ooak DataSamsung, GoPuff, YC | Enterprise Workflows | |
| PhinityNVIDIA, AWS, Stanford | Chip Design | |
| Rise Data LabsWharton | Enterprise Workflows |
No entries match. Clear the filter above to see all 80.
AI Benchmarks Catalog
Dedicated view ↗| Benchmark | Domains | Source |
|---|---|---|
| SWE-bench Pro | Multi-DomainCoding | |
| Terminal-Bench 2.0State-of-the-art set of difficult terminal-based tasks | Coding | |
| Terminal-Bench 2.1State-of-the-art set of difficult terminal-based tasks | Coding | |
| WebArenaBrowserGym leaderboard slice for WebArena, evaluating autonomous web agents across realistic browser tasks. | Agentic | |
| OSWorldA computer-use benchmark for GUI task completion across the broader OSWorld task suite. | Agentic | |
| SWE-benchVals AI’s independent run of the public SWE-bench issue set, reported by human time-to-fix bucket. Vals… | Coding | |
| Terminal-Bench / Terminal-Bench 2.0Terminal-Bench / Terminal-Bench 2.0 (arXiv:2601.11868) | CodingComputer UseEnterprise Workflows | |
| Senior SWE-Bench | Multi-DomainCoding | |
| OSWorld-Verified | Computer UseEnterprise Workflows | |
| FinanceQAFinanceQA: A Benchmark for Evaluating Financial Analysis Capabilities in LLMs (arXiv:2501.18062;… | CodingComputer UseEnterprise Workflows | |
| IDE-Bench | CodingComputer UseEnterprise Workflows | |
| EnterpriseBench CoreCraftEnterpriseBench CoreCraft: Training Generalizable Agents on High-Fidelity RL Environments - arXiv:2602.16179… | Enterprise WorkflowsLong-Horizon | |
| APEX | CodingEnterprise WorkflowsLong-Horizon | |
| APEX-Agents | CodingEnterprise WorkflowsLong-Horizon | |
| APEX-SWE | CodingEnterprise WorkflowsLong-Horizon | |
| AIME 2025All 30 problems from the 2025 American Invitational Mathematics Examination, testing olympiad-level… | Math | |
| BrowseCompBrowseComp: Evaluates autonomous agent performance on multi-step tasks requiring planning, state tracking,… | Agentic | |
| FHIR-AgentBenchFHIR-AgentBench evaluates LLM agents on realistic interoperable EHR question answering over HL7 FHIR… | Healthcare | |
| FrontierMath 2025-02-28 PrivatePrivate FrontierMath research-level mathematics benchmark snapshot. | Math | |
| FrontierMath Tier 4 2025-07-01 PrivatePrivate Tier 4 FrontierMath problems at research-level mathematical difficulty. | Math | |
| Global MMLUGlobal MMLU multilingual knowledge-and-reasoning evaluation reported in Anthropic's Claude Opus 4.8 system… | Multilingual | |
| GPQA DiamondThe hardest GPQA subset of graduate-level science questions in biology, chemistry, and physics. | Reasoning | |
| GSM8KGrade-school math word-problem benchmark for evaluating multi-step arithmetic and reasoning performance. | Math | |
| Humanity's Last ExamFrontier-level benchmark with expert-vetted closed-ended questions across mathematics, sciences, and… | Knowledge | |
| LiveCodeBench-PlusThe rapid progress of frontier large language models has led to widespread benchmark saturation, limiting the… | Coding | |
| Long-Horizon Terminal-BenchLong-Horizon Terminal-Bench evaluates whether agents can sustain progress on 46 reproducible terminal tasks… | Agentic | |
| MMLU-ProEnhanced MMLU benchmark with graduate-level questions across 14 subject areas and ten answer options. | Knowledge | |
| MultiMedia-TerminalBenchTerminals provide a powerful interface for AI agents by exposing diverse tools for automating complex… | Agentic | |
| OSWorld 2.0Anthropic first-attempt evaluation on the OSWorld authors August 2026 task release, preserving both partial… | Computer Use | |
| OTIS Mock AIME 2024-2025Competition-level math problems from OTIS Mock AIME evaluating olympiad-level math. | Math | |
| SimpleQA Net ScoreSimpleQA closed-book factuality results reported by Anthropic as net score: correct responses minus incorrect… | Knowledge | |
| SWE-bench DockerSWE-bench Docker: Evaluates software-engineering agents on realistic issue resolution, repository navigation,… | Coding | |
| SWE-bench ExtraSWE-bench Extra: Evaluates software-engineering agents on realistic issue resolution, repository navigation,… | Coding | |
| SWE-bench JavaScriptSWE-bench JavaScript: Evaluates software-engineering agents on realistic issue resolution, repository… | Coding | |
| SWE-bench VerifiedBash-only variant of SWE-bench Verified for real-world GitHub issue resolution. | Coding | |
| Terminal BenchTerminal and command-line interaction tasks for evaluating agent performance. | Agentic | |
| Terminal-Bench HardArtificial Analysis Terminal-Bench hard subset for terminal-based software engineering, system administration,… | Agentic | |
| AIMEChallenging national math exam given to top high-school students | Math | |
| LiveCodeBenchOur Implementation of the LiveCodeBench benchmark | Coding | |
| Hy-BrowseComp-Pro2Tencent internal agentic-search evaluation reported in the HY4 preview benchmark appendix; distinct from… | Search | |
| SWE-bench Verified MiniHAL's cost-aware agent leaderboard for the SWE-bench Verified Mini software engineering subset. | Coding | |
| TAU-bench AirlineHAL's standardized, cost-aware agent leaderboard for TAU-bench Airline customer-service tasks. | Agentic | |
| HELM GSM8KHELM GSM8K: Measures mathematical reasoning, symbolic problem solving, proof construction, or… | Math | |
| DeepSWE | Coding | |
| CoreCraft | Multi-Domain | |
| FinanceBench | Multi-DomainCoding | |
| SWE-Marathon | Coding | |
| SWE | CodingML | |
| Terminal-Bench 3.0A continuously maintained benchmark for difficult computer work, including coding, deep learning, finance,… | Agentic | |
| Terminal-Bench 4.0The current Terminal-Bench release measures difficult computer work after recalibrating task resources, fixing… | Agentic | |
| Terminal-Bench-Science 0.1A benchmark of AI agents completing expert-curated research workflows across the life, physical, Earth,… | Agentic | |
| HLE w/ toolsTool-augmented Humanity's Last Exam scores reported in DeepSeek-V4 thinking-mode evaluations. | Agentic | |
| AA Terminal-Bench 4.0An independently evaluated Terminal-Bench v4.0 result from Artificial Analysis, one of the ten components of… | Agentic | |
| BrowseComp-VLA vision-language browsing benchmark for multimodal web research and tool-use workflows. | Agentic | |
| TAU-benchOriginal TAU-bench evaluates a model-driven agent in simulated airline and retail customer-service… | Agentic | |
| WebArena-VerifiedWebArena-Verified is an audited release of the WebArena browser-agent benchmark. It rechecks task… | Agentic | |
| AA LiveCodeBenchAn independently evaluated LiveCodeBench result from Artificial Analysis. | Coding | |
| AA Terminal-Bench 2.1An independently evaluated Terminal-Bench v2.1 result from Artificial Analysis. | Coding | |
| HumanEvalA set of 164 handwritten Python function-generation problems. HumanEval is useful as a historical floor check,… | Coding | |
| CodeforcesCompetitive-programming rating reported for DeepSeek-V4 thinking-mode evaluations. | Coding | |
| LiveCodeBench v6LiveCodeBench v6 is a named release slice used in provider comparison tables. Keeping it separate prevents v6… | Coding | |
| LiveCodeBench v5LiveCodeBench v5 is a named release and date-window slice. BenchLM keeps explicitly labeled v5 rows outside… | Coding | |
| LiveCodeBench Pass@1-COTThis lane contains DeepSeek's LiveCodeBench Pass@1-COT results. The explicit metric and prompting label keeps… | Coding | |
| LiveCodeBench ProA harder competitive-programming benchmark family built from Codeforces, ICPC, and IOI problems, with… | Coding | |
| Multi-SWE BenchA multi-language software-engineering benchmark that measures repository-level bug fixing and implementation… | Coding | |
| ARC-AGI-1ARC Prize fluid-intelligence benchmark using novel visual grid transformations. | Reasoning | |
| ARC-AGI-2A benchmark measuring fluid intelligence and novel abstract reasoning through visual grid puzzles. Models must… | Reasoning | |
| ARC-AGI-3An interactive successor to ARC-AGI-2 that evaluates whether an AI agent can learn unfamiliar task mechanics… | Reasoning | |
| SWE-bench MultimodalA multimodal variant of SWE-bench that adds visual context (screenshots, design mockups) to software… | Multimodal | |
| AA MMLU-ProAn independently evaluated MMLU-Pro result from Artificial Analysis. | Knowledge | |
| MMLUA comprehensive multiple-choice question answering test covering 57 tasks including elementary mathematics, US… | Knowledge | |
| GPQAA challenging dataset of 448 multiple-choice questions written by domain experts in biology, physics, and… | Knowledge | |
| GPQA-DA display-only GPQA Diamond reference from provider comparison charts. | Knowledge | |
| HLEAn expert-authored benchmark designed to probe frontier knowledge and reasoning. BenchLM keeps protocol… | Knowledge | |
| HLE-VerifiedA verified and revised Humanity's Last Exam set that removes uncertain items and repairs fixable questions… | Knowledge | |
| AA-GPQA DiamondA display-only Artificial Analysis GPQA Diamond score. | Knowledge | |
| AA-HLEA display-only Artificial Analysis Humanity's Last Exam score. | Knowledge | |
| SimpleQAA benchmark that evaluates the ability of language models to answer short, fact-seeking questions accurately.… | Knowledge | |
| Chinese-SimpleQAA Chinese short-form factuality benchmark reported by DeepSeek for V4 model evaluations. | Knowledge | |
| HLE w/o toolsTool-free variant of Humanity's Last Exam that isolates a model's raw frontier reasoning. | Knowledge |
No entries match. Clear the filter above to see all 80.
Model rankings
Dedicated view ↗| Rank | Model | Score | Creator | Source | |
|---|---|---|---|---|---|
| 1 | 84.74 | Anthropic | |||
| 2 | 82.81 | OpenAI | |||
| 3 | 81.87 | Anthropic | |||
| 4 | 81.41 | Anthropic | |||
| 5 | 80.46 | OpenAI | |||
| 6 | 75.89 | ||||
| 7 | 74.43 | Moonshot AI | |||
| 8 | 73.17 | Alibaba | |||
| 9 | 72.17 | Anthropic | |||
| 10 | 71.86 | OpenAI | |||
| 11 | 70.88 | OpenAI | |||
| 12 | 70.60 | Anthropic | |||
| 13 | 70.14 | ||||
| 14 | 69.74 | xAI | |||
| 15 | 69.68 | Anthropic | |||
| 16 | 69.21 | Anthropic | |||
| 17 | 68.52 | ||||
| 18 | 68.42 | ||||
| 19 | 68.02 | ||||
| 20 | 67.99 | xAI | |||
| 21 | 67.27 | xAI | |||
| 22 | 66.93 | Z.AI | |||
| 23 | 66.87 | Alibaba | |||
| 24 | 66.68 | Z.AI | |||
| 25 | 66.36 | ||||
| 26 | 65.60 | Moonshot AI | |||
| 27 | 65.50 | Moonshot AI | |||
| 28 | 65.32 | Ornith AI | |||
| 29 | 64.86 | OpenAI | |||
| 30 | 64.29 | Alibaba | |||
| 31 | 64.14 | Dots Studio | |||
| 32 | 64.03 | OpenAI | |||
| 33 | 63.94 | Alibaba | |||
| 34 | 63.69 | DeepSeek | |||
| 35 | 63.26 | Z.AI | |||
| 36 | 63.18 | OpenAI | |||
| 37 | 62.86 | Anthropic | |||
| 38 | 62.51 | Xiaomi | |||
| 39 | 61.94 | ||||
| 40 | 61.60 | Alibaba | |||
| 41 | 61.47 | Z.AI | |||
| 42 | 61.25 | MiniMax | |||
| 43 | 60.89 | Tencent | |||
| 44 | 60.79 | Tencent | |||
| 45 | 60.38 | Thinking Machines Lab | |||
| 46 | 60.36 | Alibaba | |||
| 47 | 59.73 | xAI | |||
| 48 | 58.78 | ||||
| 49 | 58.42 | Anthropic | |||
| 50 | 57.99 | xAI |
No entries match. Clear the filter above to see all 50.