AI Benchmarks Catalog
Pick a benchmark below. Each row opens a page with publisher, domains, source, and related environments.
80 of 3,144 tracked entries · snapshot 2026-09-21 · type to filter, click a header to sort
| Benchmark | Domains | Source |
|---|---|---|
| SWE-bench Pro | Multi-DomainCoding | |
| Terminal-Bench 2.0State-of-the-art set of difficult terminal-based tasks | Coding | |
| Terminal-Bench 2.1State-of-the-art set of difficult terminal-based tasks | Coding | |
| WebArenaBrowserGym leaderboard slice for WebArena, evaluating autonomous web agents across realistic browser tasks. | Agentic | |
| OSWorldA computer-use benchmark for GUI task completion across the broader OSWorld task suite. | Agentic | |
| SWE-benchVals AI’s independent run of the public SWE-bench issue set, reported by human time-to-fix bucket. Vals… | Coding | |
| Terminal-Bench / Terminal-Bench 2.0Terminal-Bench / Terminal-Bench 2.0 (arXiv:2601.11868) | CodingComputer UseEnterprise Workflows | |
| Senior SWE-Bench | Multi-DomainCoding | |
| OSWorld-Verified | Computer UseEnterprise Workflows | |
| FinanceQAFinanceQA: A Benchmark for Evaluating Financial Analysis Capabilities in LLMs (arXiv:2501.18062;… | CodingComputer UseEnterprise Workflows | |
| IDE-Bench | CodingComputer UseEnterprise Workflows | |
| EnterpriseBench CoreCraftEnterpriseBench CoreCraft: Training Generalizable Agents on High-Fidelity RL Environments - arXiv:2602.16179… | Enterprise WorkflowsLong-Horizon | |
| APEX | CodingEnterprise WorkflowsLong-Horizon | |
| APEX-Agents | CodingEnterprise WorkflowsLong-Horizon | |
| APEX-SWE | CodingEnterprise WorkflowsLong-Horizon | |
| AIME 2025All 30 problems from the 2025 American Invitational Mathematics Examination, testing olympiad-level… | Math | |
| BrowseCompBrowseComp: Evaluates autonomous agent performance on multi-step tasks requiring planning, state tracking,… | Agentic | |
| FHIR-AgentBenchFHIR-AgentBench evaluates LLM agents on realistic interoperable EHR question answering over HL7 FHIR… | Healthcare | |
| FrontierMath 2025-02-28 PrivatePrivate FrontierMath research-level mathematics benchmark snapshot. | Math | |
| FrontierMath Tier 4 2025-07-01 PrivatePrivate Tier 4 FrontierMath problems at research-level mathematical difficulty. | Math | |
| Global MMLUGlobal MMLU multilingual knowledge-and-reasoning evaluation reported in Anthropic's Claude Opus 4.8 system… | Multilingual | |
| GPQA DiamondThe hardest GPQA subset of graduate-level science questions in biology, chemistry, and physics. | Reasoning | |
| GSM8KGrade-school math word-problem benchmark for evaluating multi-step arithmetic and reasoning performance. | Math | |
| Humanity's Last ExamFrontier-level benchmark with expert-vetted closed-ended questions across mathematics, sciences, and… | Knowledge | |
| LiveCodeBench-PlusThe rapid progress of frontier large language models has led to widespread benchmark saturation, limiting the… | Coding | |
| Long-Horizon Terminal-BenchLong-Horizon Terminal-Bench evaluates whether agents can sustain progress on 46 reproducible terminal tasks… | Agentic | |
| MMLU-ProEnhanced MMLU benchmark with graduate-level questions across 14 subject areas and ten answer options. | Knowledge | |
| MultiMedia-TerminalBenchTerminals provide a powerful interface for AI agents by exposing diverse tools for automating complex… | Agentic | |
| OSWorld 2.0Anthropic first-attempt evaluation on the OSWorld authors August 2026 task release, preserving both partial… | Computer Use | |
| OTIS Mock AIME 2024-2025Competition-level math problems from OTIS Mock AIME evaluating olympiad-level math. | Math | |
| SimpleQA Net ScoreSimpleQA closed-book factuality results reported by Anthropic as net score: correct responses minus incorrect… | Knowledge | |
| SWE-bench DockerSWE-bench Docker: Evaluates software-engineering agents on realistic issue resolution, repository navigation,… | Coding | |
| SWE-bench ExtraSWE-bench Extra: Evaluates software-engineering agents on realistic issue resolution, repository navigation,… | Coding | |
| SWE-bench JavaScriptSWE-bench JavaScript: Evaluates software-engineering agents on realistic issue resolution, repository… | Coding | |
| SWE-bench VerifiedBash-only variant of SWE-bench Verified for real-world GitHub issue resolution. | Coding | |
| Terminal BenchTerminal and command-line interaction tasks for evaluating agent performance. | Agentic | |
| Terminal-Bench HardArtificial Analysis Terminal-Bench hard subset for terminal-based software engineering, system administration,… | Agentic | |
| AIMEChallenging national math exam given to top high-school students | Math | |
| LiveCodeBenchOur Implementation of the LiveCodeBench benchmark | Coding | |
| Hy-BrowseComp-Pro2Tencent internal agentic-search evaluation reported in the HY4 preview benchmark appendix; distinct from… | Search | |
| SWE-bench Verified MiniHAL's cost-aware agent leaderboard for the SWE-bench Verified Mini software engineering subset. | Coding | |
| TAU-bench AirlineHAL's standardized, cost-aware agent leaderboard for TAU-bench Airline customer-service tasks. | Agentic | |
| HELM GSM8KHELM GSM8K: Measures mathematical reasoning, symbolic problem solving, proof construction, or… | Math | |
| DeepSWE | Coding | |
| CoreCraft | Multi-Domain | |
| FinanceBench | Multi-DomainCoding | |
| SWE-Marathon | Coding | |
| SWE | CodingML | |
| Terminal-Bench 3.0A continuously maintained benchmark for difficult computer work, including coding, deep learning, finance,… | Agentic | |
| Terminal-Bench 4.0The current Terminal-Bench release measures difficult computer work after recalibrating task resources, fixing… | Agentic | |
| Terminal-Bench-Science 0.1A benchmark of AI agents completing expert-curated research workflows across the life, physical, Earth,… | Agentic | |
| HLE w/ toolsTool-augmented Humanity's Last Exam scores reported in DeepSeek-V4 thinking-mode evaluations. | Agentic | |
| AA Terminal-Bench 4.0An independently evaluated Terminal-Bench v4.0 result from Artificial Analysis, one of the ten components of… | Agentic | |
| BrowseComp-VLA vision-language browsing benchmark for multimodal web research and tool-use workflows. | Agentic | |
| TAU-benchOriginal TAU-bench evaluates a model-driven agent in simulated airline and retail customer-service… | Agentic | |
| WebArena-VerifiedWebArena-Verified is an audited release of the WebArena browser-agent benchmark. It rechecks task… | Agentic | |
| AA LiveCodeBenchAn independently evaluated LiveCodeBench result from Artificial Analysis. | Coding | |
| AA Terminal-Bench 2.1An independently evaluated Terminal-Bench v2.1 result from Artificial Analysis. | Coding | |
| HumanEvalA set of 164 handwritten Python function-generation problems. HumanEval is useful as a historical floor check,… | Coding | |
| CodeforcesCompetitive-programming rating reported for DeepSeek-V4 thinking-mode evaluations. | Coding | |
| LiveCodeBench v6LiveCodeBench v6 is a named release slice used in provider comparison tables. Keeping it separate prevents v6… | Coding | |
| LiveCodeBench v5LiveCodeBench v5 is a named release and date-window slice. BenchLM keeps explicitly labeled v5 rows outside… | Coding | |
| LiveCodeBench Pass@1-COTThis lane contains DeepSeek's LiveCodeBench Pass@1-COT results. The explicit metric and prompting label keeps… | Coding | |
| LiveCodeBench ProA harder competitive-programming benchmark family built from Codeforces, ICPC, and IOI problems, with… | Coding | |
| Multi-SWE BenchA multi-language software-engineering benchmark that measures repository-level bug fixing and implementation… | Coding | |
| ARC-AGI-1ARC Prize fluid-intelligence benchmark using novel visual grid transformations. | Reasoning | |
| ARC-AGI-2A benchmark measuring fluid intelligence and novel abstract reasoning through visual grid puzzles. Models must… | Reasoning | |
| ARC-AGI-3An interactive successor to ARC-AGI-2 that evaluates whether an AI agent can learn unfamiliar task mechanics… | Reasoning | |
| SWE-bench MultimodalA multimodal variant of SWE-bench that adds visual context (screenshots, design mockups) to software… | Multimodal | |
| AA MMLU-ProAn independently evaluated MMLU-Pro result from Artificial Analysis. | Knowledge | |
| MMLUA comprehensive multiple-choice question answering test covering 57 tasks including elementary mathematics, US… | Knowledge | |
| GPQAA challenging dataset of 448 multiple-choice questions written by domain experts in biology, physics, and… | Knowledge | |
| GPQA-DA display-only GPQA Diamond reference from provider comparison charts. | Knowledge | |
| HLEAn expert-authored benchmark designed to probe frontier knowledge and reasoning. BenchLM keeps protocol… | Knowledge | |
| HLE-VerifiedA verified and revised Humanity's Last Exam set that removes uncertain items and repairs fixable questions… | Knowledge | |
| AA-GPQA DiamondA display-only Artificial Analysis GPQA Diamond score. | Knowledge | |
| AA-HLEA display-only Artificial Analysis Humanity's Last Exam score. | Knowledge | |
| SimpleQAA benchmark that evaluates the ability of language models to answer short, fact-seeking questions accurately.… | Knowledge | |
| Chinese-SimpleQAA Chinese short-form factuality benchmark reported by DeepSeek for V4 model evaluations. | Knowledge | |
| HLE w/o toolsTool-free variant of Humanity's Last Exam that isolates a model's raw frontier reasoning. | Knowledge |
No entries match. Clear the filter above to see all 80.