findings 5 min read

Coding agents are tied on correctness and 23 points apart on judgment. Surge's sudo L7 measures the staff-engineer gap

45% top task success on 60 tasks; 1.1-point gap on correctness against 23.2 on thought partnership; about $1 to $10 per task

sudo L7: top models 1.1 points apart on functional correctness and 23.2 points apart on thought partnership

The top coding agents now write equally correct code and differ by more than 20 points on whether to write it at all. Surge AI’s sudo L7, published October 8, 2026, is a benchmark of 60 tasks written by staff-level-or-above engineers, most of them drawn from private production repositories, graded on eight dimensions of engineering behavior. Claude Opus 5.5 and Sonnet 5.5 lead at 45% task success, GPT-6 Astra scores 32.2%, GPT-6.1 Sol 31.7%, and Claude Fable 5.1 27.2%. On functional correctness the top four models sit within 3 points of one another, 82.4% to 79.6%; on thought partnership the Claude pair leads the GPT pair by 23.2 points, and on communication by 22.7. The cost frontier runs from Sol at about $1 per task to Opus 5.5 at about $10. The benchmark starts where 1,700 coding tasks left off: tests can be passed by an agent that should have declined the task, and sudo L7 is built to score the declining.

Key Takeaways

  • 60 tasks across backend, frontend, security, reliability, and architecture, in Python, TypeScript, JavaScript, Ruby, and other languages, each written by a staff-plus engineer and each starting from a place a frontier model failed.
  • Eight graded dimensions: functional correctness, engineering craft, verification, thought partnership, communication, truthful reporting, persistence, and practical judgment. Grader models score expert-written criteria using unit tests and static analysis as evidence.
  • The GPT models lead on truthful reporting (Astra 89.4%, Sol 85.6%, against Opus 5.5 at 75.3%); the Claude models lead on thought partnership, communication, and verification.

What does sudo L7 measure that SWE-bench does not?

Surge’s framing is a promotion ladder. “Coding agents are getting very good at being L3s.” The benchmark asks whether an agent can notice risks, question a request, choose a sound approach, verify the dangerous cases, and report honestly, which Surge separates from the inner loop of writing code that passes tests. Every task was authored by a staff-level-or-above engineer, with contributors from YC, NASA JPL, Ramp, Google, Meta, Cloudera, and GitHub, and “Every task begins with a place where a frontier model meaningfully failed.” The repositories are mostly private production code, one of them a money-management application of about 250,000 lines.

The contrast with existing coding benchmarks is explicit. SWE-bench asks whether an agent can fix a real issue without breaking existing behavior; sudo L7 asks what happens after a correct patch, such as whether deploying it could cause harm. Terminal-Bench checks whether a specified task gets done; sudo L7 leaves open whether the specified approach was right. Two of Surge’s examples make the distinction concrete. GPT-6 Astra followed a request to log failed webhook payloads and return HTTP 200 without warning that this could permanently lose events and expose financial data. Opus 5.5 built a new donation-splitting system for an app that already had one, leaving two overlapping code paths behind 267 passing tests.

How do the models score?

Surge's sudo L7 leaderboard: Claude Opus 5.5 and Sonnet 5.5 at 45%, GPT-6 Astra 32.2%, GPT-6.1 Sol 31.7%, Claude Fable 5.1 27.2%, Qwen 3.8 Max 15.6%, GLM 5.3 15%, Kimi K3 13.9%, Muse Spark 1.3 12.2%, Hy4 Preview 11.1%, Grok 4.7 and Gemini 3.8 Flash 10.6%
Task success at launch. Twelve models, a 45% ceiling, and a gap of nearly 13 points between the Claude pair and the GPT pair.

Chart: Surge AI, October 8, 2026 · source

Task success is the headline, and the per-dimension criterion pass rates are the finding. On functional correctness the top four are nearly tied: GPT-6.1 Sol 82.4%, Opus 5.5 82.1%, GPT-6 Astra 81.5%, Sonnet 5.5 79.6%. On thought partnership, whether the agent questioned or improved the request, Opus 5.5 scores 76.1% and Sonnet 5.5 74.8% against Sol’s 53.0% and Astra’s 51.6%. On communication the split is 66.4% and 72.4% against 44.1% and 49.4%. The GPT models win one dimension outright, truthful reporting, at 89.4% and 85.6%.

Surge's chart of criterion pass rates for the top four models: functional correctness within 3 points (82.4% to 79.6%), thought partnership with Claude ahead by 23.2 points, communication with Claude ahead by 22.7 points
Equal on correctness, unequal on judgment. The same four models, three dimensions, and the gaps the task-success number hides.

Chart: Surge AI, October 8, 2026 · source

sudo L7 gaps between the Claude pair and the GPT pair: 1.1 points on functional correctness, 23.2 on thought partnership, 22.7 on communication

Cost tracks score only at the top. Surge’s cost-performance chart plots Sol at 29.8% for about $1 per task, Astra 2.3 points higher at roughly five times the cost, and Opus 5.5 at 45.0% for about $10; Sonnet 5.5 matches Opus at about $14. Below the frontier, Fable 5.1 spends about $9 for 27.2%.

What does a judgment benchmark mean for environment builders?

Three things. First, the grading design: expert-written criteria per task, scored by grader models with executable evidence, is the semantic-reward pattern our checklist comparison found vendors adopting for open-ended work, and it carries the judge-reliability risk that comes with it. Surge mitigates by anchoring criteria to tests and static analysis where it can. Second, the reward target: 267 passing tests on redundant code is the clearest published case of an outcome check that is technically satisfied and professionally wrong, and a training signal built only on tests would have rewarded it. Third, the market: Surge now has coding tasks that lifted a model on five benchmarks, and a benchmark on which that same style of task scores 45%. “Generating code is becoming cheap. Engineering judgment is not.”

What this means

The frontier on coding correctness has closed to within a few points across labs, and the open gap is judgment about what to build. Environments that score only whether tests pass will train agents that build the wrong thing correctly; the next generation of coding tasks needs criteria for the decision, not only the diff.

FAQ

What is sudo L7?

Surge AI’s benchmark for staff-level software engineering: 60 tasks authored by staff-plus engineers, mostly in private production repositories, graded on eight dimensions of engineering behavior by grader models using expert-written criteria and executable checks as evidence. Opus 5.5 and Sonnet 5.5 lead at 45%.

Why do GPT and Claude models tie on correctness but not on judgment?

Surge’s per-dimension data show the top four models within 3 points on functional correctness, while the Claude pair leads by 23.2 points on thought partnership and 22.7 on communication, and the GPT pair leads on truthful reporting. The benchmark does not explain the cause; it measures the behaviors separately so the difference is visible.

How much does a sudo L7 task cost to run?

Surge’s cost chart puts GPT-6.1 Sol at about $1 per task, GPT-6 Astra at about $5, Claude Opus 5.5 at about $10, and Claude Sonnet 5.5 at about $14, with the frontier flattening above $10.