findings Updated 6 min read

Mercor's APEX-Accounting benchmark measures month-end close and leaves audit out of scope

12,840 subledger rows in one reference audit case

APEX-Accounting grades month-end close; audit work is graded on whether each of 12,840 subledger rows traces to source evidence

Mercor and Ramp released APEX-Accounting on July 31, 2026: a benchmark for month-end close and bookkeeping workflows with 160 held-out (unpublished) tasks across 10 simulated companies, expert-authored rubrics averaging 13.7 criteria per task, and an open-source AI judge that Mercor reports at 97% agreement with expert graders. A benchmark of this kind cannot carry evidence integrity, which is what audit work is graded on, and Mercor’s own scope note says as much: APEX-Accounting does not evaluate tax, audit, consolidation, or external reporting. A reference audit case we examined makes the distinction concrete: a 12,840-row inventory subledger totaling $8.305 million, in which the graded competency is whether each entry traces to a source document and unsupported entries are escalated. Claude Fable 5 tops the APEX-Accounting leaderboard at 56.4%, and the most consistent model solved 2.6% of tasks correctly in all eight runs. Accounting tasks were among the 96.5% without definitive benchmark coverage in June; the close-and-bookkeeping part of that gap is closing, and the evidence-integrity part is not.

Key Takeaways

  • APEX-Accounting (July 31, 2026) brings the APEX method to accounting: real month-end close workflows, expert rubrics averaging 13.7 criteria per task, an AI judge at 97% agreement with experts, and an open 10-task sample. The full 160-task set stays closed.
  • What no reconciliation benchmark can carry is evidence integrity: whether every number traces to a source document and unsupported entries get escalated. Audit work is graded on that chain, and Mercor’s scope note puts audit outside the benchmark.
  • Payroll, medical billing, claims adjusting, freight dispatch, and bookkeeping on a real ledger still have no agent environment in our catalog.

What does APEX-Accounting get right?

APEX-Accounting applies the method that made APEX-Agents credible: real workflows rather than isolated calculations, and expert-authored criteria. The APEX family is documented on arXiv. The original index built expert-written tasks across banking, consulting, law, and medicine (Vidgen et al., 2025), and the agentic follow-on put long-horizon, cross-application tasks on arXiv (Vidgen et al., 2026). APEX-Agents launched with 480 tasks across 33 simulated worlds in investment banking, management consulting, and law. APEX-Accounting narrows the focus to one domain: 160 tasks across 10 simulated companies, each frozen at month-end close with its own accounts, records, and accounting software, authored by experts with prior experience at Deloitte, PwC, EY, and KPMG.

The release is partly open. The full set is closed and Mercor runs leaderboard evaluations on request; a sample of 1 world and 10 tasks is on Hugging Face, and the evaluation framework is on GitHub. The headline results: Claude Fable 5 at 56.4%, Meta’s Muse Spark 1.1 at 52.6%, and GPT-5.6 Sol at 51.5%, with the most consistent model solving 2.6% of tasks correctly in all eight runs.

It is a step past the last generation of finance evals, which were mostly question answering: FinQA tested numerical reasoning over report snippets, and FinanceBench showed retrieval-equipped GPT-4 still hallucinating on open-book questions about real filings. Workflow execution with expert criteria is a different, harder category.

What can APEX-Accounting not test?

It cannot test evidence integrity: whether every number traces to a source document. Real audit work is graded on that chain: a subledger entry, a journal voucher, a shipping receipt, a receiving log. A reconciliation task asks whether the GL (general ledger) matches the subledger, which tests arithmetic. An audit task asks whether every GL entry traces to source evidence and whether unsupported entries are escalated, which tests the competency audit is graded on. Mercor says so itself: audit is out of scope, along with tax, consolidation, and external reporting.

The research on LLM auditing finds the same split. Models can spot statement errors at reasonable rates but fail at explaining them, citing the governing accounting standards, and completing a full audit workflow (Wang et al., 2025). Models find wrong numbers; they do not yet carry the evidence chain, which is the graded competency.

Our reference audit case makes it concrete. The case is built on a 12,840-row inventory subledger totaling $8.305 million, seeded with planted discrepancies; the graded competency is whether each entry traces to a source document and unsupported entries are escalated. A reconciliation benchmark would test whether the agent matches the GL to the subledger. An evidence-integrity environment tests whether the agent identifies every unsupported entry, traces each number to its source, and escalates the control failures. Pass rates on the first do not predict the second. What that environment’s spec contains is covered separately.

Which domains are still open?

Accounting tasks were among the 96.5% of tasks without definitive benchmark coverage in the June coverage study. APEX-Accounting closes part of that gap, the month-end close and bookkeeping part. The evidence-integrity part remains open, and it is the harder part to build.

These domains had no agent environment in our catalog as of July 31, 2026: the five listed in the Deeptune post (payroll, medical billing, claims adjusting, freight dispatch, hotel property management) plus three more:

  • bookkeeping on a real ledger
  • support ticketing on real software
  • warehouse management

All of these are desk work, so they pass the physical-dependency screen; what they lack is a vendor.

“AI models can pass the CPA exam”

Mercor's APEX-Accounting leaderboard: Claude Fable 5 at 56.4%, Muse Spark 1.1 at 52.6%, GPT-5.6 Sol at 51.5%

Mercor's own launch chart, posted the day the benchmark shipped.

Mercor (@mercor) · July 31, 2026 · on X

What this means

APEX-Accounting covers month-end close and bookkeeping with a grader labs can inspect. Evidence integrity, the source-document chain audit work is graded on, still has no benchmark, and neither do payroll, claims, or freight dispatch. Those are the domains a consolidating buyer cannot reach through credential networks, which leaves them to independent vendors.

FAQ

What is evidence integrity in audit work?

Evidence integrity means every number in a workpaper traces to a source document: a subledger entry, a journal voucher, a shipping receipt. An agent that produces a numerically correct total without source-document support has not completed the audit.

What is the 96.5% coverage gap?

The June coverage study found that only 3.5% of 202 O*NET occupational tasks have definitive benchmark coverage. The remaining 96.5% are partial-only, unverified, or uncovered. Accounting tasks were in that gap, and APEX-Accounting closes the close-and-bookkeeping part of it.

What is APEX-Accounting?

APEX-Accounting is a benchmark released by Mercor and Ramp, the corporate-card and finance-automation company, on July 31, 2026. It holds 160 closed tasks across 10 simulated companies for month-end close and bookkeeping, graded by an open-source AI judge against expert rubrics, with a 10-task open sample. It is part of Mercor’s APEX benchmark family.