findings 5 min read

Can AI agents survive a 9 to 5? Surge's DAYJOB puts the best model under 25% in finance and healthcare

Best model 24.7% (healthcare) and 23.9% (finance) on 130 expert assignments; prompts average 49 to 81 words against 337 in GDPval

The best agent completes 24.7% of DAYJOB healthcare assignments and 23.9% of finance assignments

When a professional task is specified the way a manager would specify it, the best frontier agent completes under a quarter of them. Surge AI’s DAYJOB, published September 23, 2026, is a pair of benchmarks, DAYJOB: Healthcare and DAYJOB: Finance, built from 130 expert assignments on which the strongest model succeeds 24.7% and 23.9% of the time. The design change from Surge’s earlier work is the brief. A DAYJOB prompt averages 81 words in finance and 49 in healthcare, against 337 in OpenAI’s GDPval; a task ships with 25.7 input files in finance and 19.8 in healthcare, against 1.2; a human professional would need an estimated 21.6 and 19.6 hours, against a GDPval median of 4. Grading uses task-specific rubrics with a median of 57.5 criteria in finance and 47.5 in healthcare, and a pass means every criterion met. The scoping, prioritizing, and synthesis that a 337-word prompt used to do for the agent is now the agent’s job, and that is where it fails.

Key Takeaways

  • Leaderboard, pass@1 with every criterion met: Claude Opus 5.5 leads both benchmarks at 24.7% (healthcare) and 23.9% (finance); GPT-6 Astra scores 11.6% and 21.5%; Claude Fable 5.1 scores 9.6% and 19.8%. Seven more models sit between 2.0% and 14.8%.
  • The brief shrank and the workspace grew: 49 to 81 words of prompt, 20 to 26 input files, and roughly 20 hours of estimated expert time per task.
  • Documented failures: a missed $25.7M pricing error, deference to an initial diagnosis the record contradicted, and misjudged scope and priority.

What does DAYJOB change about the task?

Surge calls it an evolution of GDPval, OpenAI’s benchmark of economically valuable deliverables, and the change is in how much the task text tells the agent. “Rather than telling models exactly what to do, DAYJOB asks whether agents can figure out what needs to be done, and reason through it until the end.” Three measurable shifts follow.

DAYJOB against GDPval on three dimensions: prompt length 81 and 49 words against 337, input files 25.7 and 19.8 against 1.2, estimated human hours 21.6 and 19.6 against a median of 4

The prompts are shorter because a manager’s request is short: 81 words on average in finance, 49 in healthcare, against 337 in GDPval. The workspaces are larger because the evidence is spread across files: 25.7 input files in finance and 19.8 in healthcare, against 1.2. The horizon is longer: 21.6 and 19.6 estimated human hours, against a GDPval median of 4. The assignments were written with domain experts, among them a senior portfolio manager at BlackRock, a finance professional at HSBC, and a pediatric endocrinologist, and passed three layers of expert review. “DAYJOB is our evolution of GDPval for long-horizon, economically valuable agents.”

How do the models score?

Surge's DAYJOB leaderboards: on healthcare, Claude Opus 5.5 24.7%, GPT-6 Astra 11.6%, Claude Fable 5.1 9.6%, down to Muse Spark 1.2 at 2.0%; on finance, Opus 5.5 23.9%, GPT-6 Astra 21.5%, Fable 5.1 19.8%, down to GLM 5.3 at 6.0%
Surge's leaderboards at launch, pass@1 with every rubric criterion met. Finance is a tight race below a low ceiling; healthcare falls off a cliff after the leader.

Chart: Surge AI, September 23, 2026 · source

“The strongest models score less than 25% on both DAYJOB: Healthcare and DAYJOB: Finance.” Claude Opus 5.5 leads both at 24.7% and 23.9%. The two domains then diverge: in finance, GPT-6 Astra (21.5%) and Claude Fable 5.1 (19.8%) sit close behind, while in healthcare the same two models drop to 11.6% and 9.6%, and the tenth-ranked model completes 2.0% of assignments. The ceiling matches Surge’s HANDBOOK.md result from June, where no model cleared 25% on strict pass@1, for a different reason. HANDBOOK.md failed agents on policy compliance inside a well-specified task; DAYJOB fails them on deciding what the task is.

Our coverage study found that existing benchmarks test isolated calculations and document operations rather than occupational workflows. DAYJOB is a benchmark built to close that gap, and the first scores say the gap is real.

What does the failure pattern say about environment design?

Surge’s examples are specific. An agent identified related pricing issues and missed a $25.7M error. An agent deferred to an initial diagnosis despite clinical findings that called for further evaluation. Agents failed to synthesize across the two dozen files in a workspace, and failed to decide what a 60-word request was asking for. These are not tool failures; they are judgment failures that a 337-word brief used to pre-empt.

For environment builders, two design choices follow. First, the brief is a scored dimension: an environment that hands the agent a complete specification measures execution, and one that hands it a manager’s request measures the job. Second, a rubric with a median of 57 criteria is a verifier design with known hazards. Multi-component rubrics leak reward through format, and a rubric judge has to be recalibrated when agents learn to game it. DAYJOB’s all-criteria pass rule limits the first hazard; the second is the cost of grading open-ended deliverables at all.

What this means

The best agent completes about one professional assignment in four when nobody scopes the work for it. For environment vendors that is the specification to build against: short briefs, large workspaces, long horizons, and a verifier that can tell a finished deliverable from a plausible one.

FAQ

What does pass@1 mean on DAYJOB?

One attempt per task, and the attempt counts as a pass only if every criterion in the task’s rubric is met. With a median of 57.5 criteria in finance and 47.5 in healthcare, a deliverable that gets most things right still fails.

How is DAYJOB different from GDPval?

GDPval gives the agent a detailed brief (337 words on average) and about one input file per task. DAYJOB gives it a manager’s request of 49 to 81 words, a workspace of 20 to 26 files, and a task that would take a professional about 20 hours. Surge built it with the same aim, economically valuable work, and more of the scoping left to the agent.

Which model leads DAYJOB?

Claude Opus 5.5, at 24.7% on healthcare and 23.9% on finance at launch on September 23, 2026. GPT-6 Astra and Claude Fable 5.1 are second and third on both, with a much larger drop behind the leader in healthcare than in finance.