findings 5 min read

AI now beats junior accountants on APEX-Accounting tasks. Mercor's study says what it leaves out

100% vs 0 to 90% of criteria met; $0.21 vs $10.35 per criterion (49x); 12 CPAs, 23 human attempts, 20 model attempts

Accountants met 0 to 90% of criteria with a 37% average; Claude Opus 5 met 100% on every attempt

Frontier models now outscore junior accountants on Mercor’s timed accounting tasks, and the study that shows it is careful about what the tasks measure. In a human baseline study published October 1, 2026, Mercor researcher Aden Barton had 12 junior accountants, all licensed CPAs averaging 5.5 years of experience, work four month-end close scenarios from APEX-Accounting under a three-hour limit. Across 23 attempts they met 0% to about 90% of the rubric criteria in 30 to 180 minutes; Claude Opus 5 met 100% on all 20 of its attempts in under ten minutes, at $0.21 per criterion against $10.35 for the accountants, 49 times less. Eighteen months earlier the best model scored below the accountants’ 37% average. Mercor’s own caveat carries the finding: the tasks test detail-oriented instruction following, the part of accounting models are best at, with no client, no coworkers, and no accumulated context. What the study measures is a benchmark reaching its ceiling.

Key Takeaways

  • Accountants without AI: 23 attempts across 4 tasks, 0% to about 90% of criteria met, 30 to 180 minutes each. Claude Opus 5: 20 attempts, 100% on every one, under 10 minutes.
  • Cost per criterion met: $0.21 for Opus 5 at list token prices with no caching discount, which Mercor calls an upper bound, against $10.35 for accountant time at the US median accountant wage.
  • Model progress on the same four tasks: GPT-4o near 4% in mid-2024, o3 about 55%, GPT-5 about 69%, Opus 4.8 about 82%, Opus 5 at 100%. The accountant average is 37%.

What did Mercor measure, and how?

The four tasks are realistic month-end close scenarios: navigate a company’s working files, perform the calculations, and deliver results in tabular form, graded against a rubric of required outputs. The accountants worked unassisted with a three-hour limit per task; the model ran the same tasks 20 times in total. “Frontier AI models are now faster and more accurate than junior accountants, even the best one in our study.”

Mercor's dot plot: 23 accountant attempts spread from 0% to about 90% of rubric criteria met and 30 to 180 minutes, while 20 Claude Opus 5 attempts sit at 100% and under 10 minutes
Every dot is one attempt, pooled across the four tasks. The accountants' best attempt is near 90%; the model's worst is 100%.

Chart: Mercor, October 1, 2026 · source

The cost comparison divides total cost by the average number of criteria met. The model’s cost is input and output tokens at list prices with no caching discount, so $0.21 is a ceiling. The accountants’ cost uses the Bureau of Labor Statistics median accountant wage rather than what the study paid them: $10.35 per criterion, 49 times more. “Models are also less expensive: comparing cost per task criterion, frontier models are more than an order of magnitude cheaper than humans.”

How fast did models cross the human line?

APEX-Accounting human baseline: accountants met 0 to 90% of criteria in 30 to 180 minutes at $10.35 per criterion; Claude Opus 5 met 100% in under 10 minutes at $0.21

Mercor ran every model it tested, budget and open-weight variants included, on the same four tasks and plotted accuracy against release date. The crossing happened in 2025. “Just eighteen months ago, the best AI models fell short of the average accountant’s ~37% score. Today, models ace those same tasks.” GPT-4o sat near 4% in mid-2024, OpenAI’s o3 near 55% in spring 2025, GPT-5 near 69% that summer, Opus 4.8 near 82% in mid-2026, and Opus 5 at 100%. Several recent releases still sit below the accountant line, Qwen3.5-122B among them at about 34%, which is why the line matters as a baseline and not as a date.

Mercor's scatter of model accuracy on the four accounting tasks by release date, from GPT-4o near 4% in mid-2024 to Opus 5 at 100% in 2026, with the 37% accountant average drawn as a horizontal band
Two years of releases against one human baseline. The shaded band is the 95% bootstrap interval around the accountants' 37% average.

Chart: Mercor, October 1, 2026 · source

What does the study say it does not measure?

Mercor is explicit. “These results do not mean accountants are replaceable, since our tasks ended up testing what AI is best at.” The tasks are medium-length, detail-oriented, and fully specified; real accounting includes client communication, unstructured requests, coworkers, and the context a person accumulates in a role, none of which the stripped-down setting contains. The study adds a second caveat about benchmarks in general: “Benchmarks aim to mirror real work, but tasks get made harder to induce model failures.”

Two of our own findings frame the result. Our July read of APEX-Accounting noted that audit, where correctness depends on evidence integrity, sits outside the benchmark’s scope by Mercor’s own note. And Surge’s DAYJOB, published nine days before this study, puts the best agent under 25% in finance when the brief is a manager’s 81-word request and the workspace holds 26 files. A task specified in full is one models now finish; a task that has to be scoped first is one they do not. For environment builders, a benchmark at 100% is a benchmark that has stopped measuring, and the fix is the one Harbor’s versioning rule prescribes: retire the saturated tasks and rerun on harder ones.

What this means

Fully specified bookkeeping tasks are solved at the frontier, at a cost two orders of magnitude below human labor. The measurable gap has moved to scoping, judgment, and evidence integrity, which is where environment vendors should now be writing tasks.

FAQ

Who were the accountants in Mercor’s study?

Twelve junior accountants, all licensed CPAs, with an average of 5.5 years of experience. They worked four realistic month-end close scenarios without AI assistance under a three-hour limit per task, producing 23 scored attempts in total.

How was cost per criterion computed?

Total cost divided by the average number of rubric criteria met. For Claude Opus 5, Mercor priced input and output tokens at list rates with no caching discount, so $0.21 is an upper bound. For accountants it used time at the US median accountant wage, giving $10.35, rather than the rate paid in the study.

Does this mean accounting work is automated?

No, by Mercor’s own account. The tasks test detail-oriented instruction following on medium-length, fully specified work, with no client, coworkers, or accumulated context. Scoping an ambiguous request and judging what matters, which other benchmarks test, remain far from saturated.