AI Benchmarks Catalog

Pick a benchmark below. Each row opens a page with publisher, domains, source, and related environments.

80 of 3,144 tracked entries · snapshot 2026-09-21 · type to filter, click a header to sort

Domain
Source

80 benchmarks shown

BenchmarkDomainsSource
SWE-bench ProMulti-DomainCodingpavlovslist
Terminal-Bench 2.0State-of-the-art set of difficult terminal-based tasksCodingbenchmarklist
Terminal-Bench 2.1State-of-the-art set of difficult terminal-based tasksCodingbenchmarklist
WebArenaBrowserGym leaderboard slice for WebArena, evaluating autonomous web agents across realistic browser tasks.Agenticbenchmarklist
OSWorldA computer-use benchmark for GUI task completion across the broader OSWorld task suite.Agenticbenchlm-benchmarks
SWE-benchVals AI’s independent run of the public SWE-bench issue set, reported by human time-to-fix bucket. Vals…Codingbenchlm-benchmarks
Terminal-Bench / Terminal-Bench 2.0Terminal-Bench / Terminal-Bench 2.0 (arXiv:2601.11868)CodingComputer UseEnterprise Workflowsrl-list
Senior SWE-BenchMulti-DomainCodingpavlovslist
OSWorld-VerifiedComputer UseEnterprise Workflowsrl-list
FinanceQAFinanceQA: A Benchmark for Evaluating Financial Analysis Capabilities in LLMs (arXiv:2501.18062;…CodingComputer UseEnterprise Workflowsrl-list
IDE-BenchCodingComputer UseEnterprise Workflowsrl-list
EnterpriseBench CoreCraftEnterpriseBench CoreCraft: Training Generalizable Agents on High-Fidelity RL Environments - arXiv:2602.16179…Enterprise WorkflowsLong-Horizonrl-list
APEXCodingEnterprise WorkflowsLong-Horizonrl-list
APEX-AgentsCodingEnterprise WorkflowsLong-Horizonrl-list
APEX-SWECodingEnterprise WorkflowsLong-Horizonrl-list
AIME 2025All 30 problems from the 2025 American Invitational Mathematics Examination, testing olympiad-level…Mathbenchmarklist
BrowseCompBrowseComp: Evaluates autonomous agent performance on multi-step tasks requiring planning, state tracking,…Agenticbenchmarklist
FHIR-AgentBenchFHIR-AgentBench evaluates LLM agents on realistic interoperable EHR question answering over HL7 FHIR…Healthcarebenchmarklist
FrontierMath 2025-02-28 PrivatePrivate FrontierMath research-level mathematics benchmark snapshot.Mathbenchmarklist
FrontierMath Tier 4 2025-07-01 PrivatePrivate Tier 4 FrontierMath problems at research-level mathematical difficulty.Mathbenchmarklist
Global MMLUGlobal MMLU multilingual knowledge-and-reasoning evaluation reported in Anthropic's Claude Opus 4.8 system…Multilingualbenchmarklist
GPQA DiamondThe hardest GPQA subset of graduate-level science questions in biology, chemistry, and physics.Reasoningbenchmarklist
GSM8KGrade-school math word-problem benchmark for evaluating multi-step arithmetic and reasoning performance.Mathbenchmarklist
Humanity's Last ExamFrontier-level benchmark with expert-vetted closed-ended questions across mathematics, sciences, and…Knowledgebenchmarklist
LiveCodeBench-PlusThe rapid progress of frontier large language models has led to widespread benchmark saturation, limiting the…Codingbenchmarklist
Long-Horizon Terminal-BenchLong-Horizon Terminal-Bench evaluates whether agents can sustain progress on 46 reproducible terminal tasks…Agenticbenchmarklist
MMLU-ProEnhanced MMLU benchmark with graduate-level questions across 14 subject areas and ten answer options.Knowledgebenchmarklist
MultiMedia-TerminalBenchTerminals provide a powerful interface for AI agents by exposing diverse tools for automating complex…Agenticbenchmarklist
OSWorld 2.0Anthropic first-attempt evaluation on the OSWorld authors August 2026 task release, preserving both partial…Computer Usebenchmarklist
OTIS Mock AIME 2024-2025Competition-level math problems from OTIS Mock AIME evaluating olympiad-level math.Mathbenchmarklist
SimpleQA Net ScoreSimpleQA closed-book factuality results reported by Anthropic as net score: correct responses minus incorrect…Knowledgebenchmarklist
SWE-bench DockerSWE-bench Docker: Evaluates software-engineering agents on realistic issue resolution, repository navigation,…Codingbenchmarklist
SWE-bench ExtraSWE-bench Extra: Evaluates software-engineering agents on realistic issue resolution, repository navigation,…Codingbenchmarklist
SWE-bench JavaScriptSWE-bench JavaScript: Evaluates software-engineering agents on realistic issue resolution, repository…Codingbenchmarklist
SWE-bench VerifiedBash-only variant of SWE-bench Verified for real-world GitHub issue resolution.Codingbenchmarklist
Terminal BenchTerminal and command-line interaction tasks for evaluating agent performance.Agenticbenchmarklist
Terminal-Bench HardArtificial Analysis Terminal-Bench hard subset for terminal-based software engineering, system administration,…Agenticbenchmarklist
AIMEChallenging national math exam given to top high-school studentsMathbenchmarklist
LiveCodeBenchOur Implementation of the LiveCodeBench benchmarkCodingbenchmarklist
Hy-BrowseComp-Pro2Tencent internal agentic-search evaluation reported in the HY4 preview benchmark appendix; distinct from…Searchbenchmarklist
SWE-bench Verified MiniHAL's cost-aware agent leaderboard for the SWE-bench Verified Mini software engineering subset.Codingbenchmarklist
TAU-bench AirlineHAL's standardized, cost-aware agent leaderboard for TAU-bench Airline customer-service tasks.Agenticbenchmarklist
HELM GSM8KHELM GSM8K: Measures mathematical reasoning, symbolic problem solving, proof construction, or…Mathbenchmarklist
DeepSWECodingpavlovslist
CoreCraftMulti-Domainpavlovslist
FinanceBenchMulti-DomainCodingpavlovslist
SWE-MarathonCodingpavlovslist
SWECodingMLpavlovslist
Terminal-Bench 3.0A continuously maintained benchmark for difficult computer work, including coding, deep learning, finance,…Agenticbenchlm-benchmarks
Terminal-Bench 4.0The current Terminal-Bench release measures difficult computer work after recalibrating task resources, fixing…Agenticbenchlm-benchmarks
Terminal-Bench-Science 0.1A benchmark of AI agents completing expert-curated research workflows across the life, physical, Earth,…Agenticbenchlm-benchmarks
HLE w/ toolsTool-augmented Humanity's Last Exam scores reported in DeepSeek-V4 thinking-mode evaluations.Agenticbenchlm-benchmarks
AA Terminal-Bench 4.0An independently evaluated Terminal-Bench v4.0 result from Artificial Analysis, one of the ten components of…Agenticbenchlm-benchmarks
BrowseComp-VLA vision-language browsing benchmark for multimodal web research and tool-use workflows.Agenticbenchlm-benchmarks
TAU-benchOriginal TAU-bench evaluates a model-driven agent in simulated airline and retail customer-service…Agenticbenchlm-benchmarks
WebArena-VerifiedWebArena-Verified is an audited release of the WebArena browser-agent benchmark. It rechecks task…Agenticbenchlm-benchmarks
AA LiveCodeBenchAn independently evaluated LiveCodeBench result from Artificial Analysis.Codingbenchlm-benchmarks
AA Terminal-Bench 2.1An independently evaluated Terminal-Bench v2.1 result from Artificial Analysis.Codingbenchlm-benchmarks
HumanEvalA set of 164 handwritten Python function-generation problems. HumanEval is useful as a historical floor check,…Codingbenchlm-benchmarks
CodeforcesCompetitive-programming rating reported for DeepSeek-V4 thinking-mode evaluations.Codingbenchlm-benchmarks
LiveCodeBench v6LiveCodeBench v6 is a named release slice used in provider comparison tables. Keeping it separate prevents v6…Codingbenchlm-benchmarks
LiveCodeBench v5LiveCodeBench v5 is a named release and date-window slice. BenchLM keeps explicitly labeled v5 rows outside…Codingbenchlm-benchmarks
LiveCodeBench Pass@1-COTThis lane contains DeepSeek's LiveCodeBench Pass@1-COT results. The explicit metric and prompting label keeps…Codingbenchlm-benchmarks
LiveCodeBench ProA harder competitive-programming benchmark family built from Codeforces, ICPC, and IOI problems, with…Codingbenchlm-benchmarks
Multi-SWE BenchA multi-language software-engineering benchmark that measures repository-level bug fixing and implementation…Codingbenchlm-benchmarks
ARC-AGI-1ARC Prize fluid-intelligence benchmark using novel visual grid transformations.Reasoningbenchlm-benchmarks
ARC-AGI-2A benchmark measuring fluid intelligence and novel abstract reasoning through visual grid puzzles. Models must…Reasoningbenchlm-benchmarks
ARC-AGI-3An interactive successor to ARC-AGI-2 that evaluates whether an AI agent can learn unfamiliar task mechanics…Reasoningbenchlm-benchmarks
SWE-bench MultimodalA multimodal variant of SWE-bench that adds visual context (screenshots, design mockups) to software…Multimodalbenchlm-benchmarks
AA MMLU-ProAn independently evaluated MMLU-Pro result from Artificial Analysis.Knowledgebenchlm-benchmarks
MMLUA comprehensive multiple-choice question answering test covering 57 tasks including elementary mathematics, US…Knowledgebenchlm-benchmarks
GPQAA challenging dataset of 448 multiple-choice questions written by domain experts in biology, physics, and…Knowledgebenchlm-benchmarks
GPQA-DA display-only GPQA Diamond reference from provider comparison charts.Knowledgebenchlm-benchmarks
HLEAn expert-authored benchmark designed to probe frontier knowledge and reasoning. BenchLM keeps protocol…Knowledgebenchlm-benchmarks
HLE-VerifiedA verified and revised Humanity's Last Exam set that removes uncertain items and repairs fixable questions…Knowledgebenchlm-benchmarks
AA-GPQA DiamondA display-only Artificial Analysis GPQA Diamond score.Knowledgebenchlm-benchmarks
AA-HLEA display-only Artificial Analysis Humanity's Last Exam score.Knowledgebenchlm-benchmarks
SimpleQAA benchmark that evaluates the ability of language models to answer short, fact-seeking questions accurately.…Knowledgebenchlm-benchmarks
Chinese-SimpleQAA Chinese short-form factuality benchmark reported by DeepSeek for V4 model evaluations.Knowledgebenchlm-benchmarks
HLE w/o toolsTool-free variant of Humanity's Last Exam that isolates a model's raw frontier reasoning.Knowledgebenchlm-benchmarks