ARC-AGI-2
AI benchmark · Reasoning
- Domains
- Reasoning
- Data source
- benchlm-benchmarks
- Snapshot
About ARC-AGI-2
A benchmark measuring fluid intelligence and novel abstract reasoning through visual grid puzzles. Models must identify patterns in input-output pairs and generate the correct output for unseen inputs. Considered the hardest public reasoning benchmark — average individual human performance is 66%.
Where ARC-AGI-2 sits
Its domains, and the nearest entries sharing them. Click any node to open its page.
Select a node to trace its connections. Zoom in for more space; drag to pan.
Next: browse all benchmarks. Entry from the RL Research daily scrape of public sources, 2026-09-21 snapshot.