APEX-Agents 1.1
AI benchmark · Coding
- Publisher
- Mercor
- Domains
- Coding Enterprise Workflows Long-Horizon
- On X
- @mercor ↗
- Snapshot
Open benchmark ↗mercor.com/apex/

Latest from Mercor on X
@mercor ↗Introducing the Mercor Research team. We’re building a world-class research organization, including @edwardjhu (first author of LoRA), @MinyangTian1 (SciCode, CritPt), Victor Barres (τ²-bench, τ-voice), @hrdkbhatnagar (PostTrainBench), @bertievidgen (APEX-Agents) and many others. As models get better, the questions get harder: → How do you measure a model on a job that AI itself is changing? → What data helps when a model fails at real work? → How much can a small amount of expert data improve a model? Mercor Research covers the full AI flywheel. APEX shows where a model struggles. Experts help us understand why and create data for those weak points. We post-train the model and measure it again. Learn more about our research mission:
— mercor (@mercor) Oct 8, 2026Claude Haiku 5.5 debuts at #7 on APEX-SWE. 55.1% Pass@1, ahead of last generation's Opus and Sonnet models. It achieves this score using 1/20th of Sonnet 5.5's input and output cost. Congrats @AnthropicAI on the launch. 𝗖𝗼𝗱𝗶𝗻𝗴 Anthropic positions Haiku 5.5 as a coding subagent next to Opus 5.5 and Sonnet 5.5. Its results on APEX-SWE suggest this model is a strong coding partner. APEX-SWE: 55.1% (#7) Integration: 63.0% (#15) Observability: 47.2% (#7) APEX-SWE covers two kinds of work. - 𝗢𝗯𝘀𝗲𝗿𝘃𝗮𝗯𝗶𝗹𝗶𝘁𝘆 asks the agent to debug a production-style system from logs and telemetry. Haiku 5.5 excels here and ranks #7, behind only six models. It still trails Opus 5.5 (69.8%) by 22.6 pts, and a typical attempt uses about 2.5M tokens, 4.5x an Integration attempt. - 𝗜𝗻𝘁𝗲𝗴𝗿𝗮𝘁𝗶𝗼𝗻 asks the agent to build and deploy systems across services. Haiku 5.5 finishes 2.5 pts behind Opus 5.5 (65.5%), but the field is bunched, and it ranks #15. Across 4 runs per task, Haiku 5.5 solved 62.5% of APEX-SWE tasks at least once. 75 of 200 tasks were never solved. It also passes last generation's Anthropic models on APEX-SWE: Opus 4.8: 43.9% Opus 4.7: 47.6% Sonnet 5: 46.4% 𝗣𝗿𝗼𝗳𝗲𝘀𝘀𝗶𝗼𝗻𝗮𝗹 𝘄𝗼𝗿𝗸 APEX-Agents is our flagship benchmark for long-horizon, agentic professional services work. It tests workflows across management consulting, corporate law, and investment banking. APEX-Agents: 56.1% (#19) Management consulting: 50.9% (#22) Corporate law: 60.0% (#26) Investment banking: 57.5% (#13) That puts it 19.4 pts behind Sonnet 5.5 (75.5%) and 17.4 pts behind Opus 5.5 (73.5%). This matches Anthropic's own advice. Sonnet and Opus are for complex agentic work, Haiku is for scoped, high-volume tasks. 𝗖𝗼𝘀𝘁 𝗮𝗻𝗱 𝗰𝗼𝗻𝘀𝗶𝘀𝘁𝗲𝗻𝗰𝘆 Haiku 5.5 lists at $0.10 input and $0.50 output per 1M tokens. That is 1/20th of Sonnet 5.5's input and output cost for prompts up to 100K tokens, which Anthropic says covers about 90% of Haiku requests. It runs about 75% cheaper than Haiku 4.5 on average, and it is the first Haiku with an adjustable effort setting. We tested all three 5.5 models at Max effort. On APEX-SWE, Haiku 5.5's output token cost was about 35x lower than Sonnet 5.5 and about 60x lower than Opus 5.5. It also finished a typical task about 3x faster. On APEX-Agents, a typical Haiku 5.5 task used more output tokens than Sonnet 5.5 and Opus 5.5 combined. It still costs about 14x to 16x less on output tokens. Across 4 runs per task on APEX-Agents, Haiku 5.5 solved 74.6% of tasks at least once, but only 37.9% on every run. See full leaderboard:
— mercor (@mercor) Oct 8, 2026Watch Mercor's Head of AI modeling @edwardjhu to learn about AI data, evaluation, and model training. Find out why "the closer we are to how AI actually gets deployed, the more realistic the benchmarks will be."
— mercor (@mercor) Oct 7, 2026
Where APEX-Agents 1.1 sits
Its domains, and the nearest entries sharing them. Click any node to open its page.
The map could not load. The domain links still work.
More Coding AI benchmarks
Coding RL environments
Next: browse all benchmarks. Entry from the RL Research daily scrape of public sources, 2026-10-10 snapshot.