2,209 agents rewrote a coding harness in Rust in two weeks. The reward was a parity check in a fresh sandbox
By Elon HuskParody pen name · View profile2,209 agents (1,981 rewrite, 228 hillclimb); 10,000+ sandboxes; 228.7B tokens; 16,758 agent messages; cold start 736.1 ms to 51.9 ms (14.2x)

The sandbox and verifier stack that reinforcement learning (RL) training runs on has been used to ship a product, and the numbers are public. Prime Intellect reported on October 9, 2026 that it rewrote Prime Agent, its coding harness, from TypeScript to Rust with a swarm of 2,209 agents working across more than 10,000 Prime Sandboxes over two weeks: 1,981 for the port and 228 for a three-day performance hillclimb, consuming 228.7B tokens from a GLM-5.3 endpoint and exchanging 16,758 agent-to-agent messages. Every task passed through a separate verifier agent that compiled the code and ran parity checks in a fresh sandbox, and the hillclimb was given no numeric targets. The result reaches input-ready in 51.9 ms from a cold start, against 736.1 ms before, with large-session memory down from 1,130 MB to 237 MB. The essay that preceded it, On the Nature of the Swarm of October 6, explains the structure: persistent agents in a mesh around shared state, with a root that writes no code.
Key Takeaways
- Scale: 2,209 agents, 10,000-plus microVM sandboxes, 228.70B tokens (192.99B for the rewrite, 35.70B for the hillclimb), 16,758 messages, two weeks. The largest source file shrank from about 15,000 lines of TypeScript to about 2,500 of Rust.
- Verification: each task had its own verifier agent compiling and running parity checks in a fresh 4-core, 8 GB sandbox; a TUI verifier compared terminal frames between builds; noise checks withheld unstable results.
- Hillclimb results: cold start 736.1 ms to 51.9 ms (14.18x), warm start 552.3 to 41.4 ms (13.34x), large-session memory 1,130 MB to 237.3 MB (4.76x), installed size 172.1 MB to 59.6 MB (2.89x), with 144 experiment records and 69 merges in three days.
How was the work verified?
Prime’s rewrite is interesting to environment builders for one design choice: the reward was objective. Each port task went to a verifier agent that compiled the Rust, ran parity checks against the TypeScript behavior, and did so in a fresh sandbox rather than the one the working agent had been modifying. A separate terminal-frame verifier compared what the two builds drew on screen. “With an objective check for every kind of parity, the agents could measure their own progress”, which is the condition under which a swarm can run without a human grading each change.
The hillclimb added a second choice worth copying. The loop ran for three days against a benchmark suite on fresh 4-core, 8 GB sandboxes, logging over 144 experiment records and merging over 69 changes, and withheld any result its noise checks flagged as unstable. They also refused to set a goal: “We deliberately gave the loop no numeric targets, since a fixed threshold tends to become a stopping point.” That is the reward design problem stated from the engineering side: a fixed target gets reached and gamed, while a measured direction keeps the search alive.
What does the swarm look like?
The structure comes from Konstantin Dunas’s essay of October 6. Its premise is that inference scaling stops when one agent fills its context window, and that compaction, summarizing and restarting, only defers the problem: “The fundamental principle behind this limitation is that every compaction is a bet.” Persistent sub-agents, which keep their context alive and can be asked follow-up questions, add roughly one window of capacity per agent, and a tree of them routes too much through the root. “The root agent is a central planner.” Dunas argues for a mesh of peers around shared state, a repository, an issue tracker, a chat, with trees allowed inside it.
The rewrite applied that shape. The root agent “wrote no product code, keeping it free to monitor every task, merge finished work” and route it, while the working agents and their verifiers ran in parallel sandboxes and talked to one another directly. At a 4-core, 8 GB sandbox priced at Prime’s list rates, about $0.18 an hour, the substrate for 10,000 concurrent sandboxes is roughly $1,800 an hour before tokens; the post gives no dollar figure for either.
What does this say about RL infrastructure?
The same three components an RL environment needs, isolated execution, a deterministic verifier, and a reward that resists gaming, built a product in two weeks with no training step at all. Prime says the stack is meant to power swarms, autonomous research, and RL; the rewrite is its demonstration. The verifier design matches the pattern the reference spec prescribes: one objective outcome check per task, run in a clean environment the agent cannot have tampered with. The hillclimb matches the pattern our analysis of self-improvement loops found missing elsewhere: a grading step that does not need a human, because parity and latency are measurable. Prime’s comparison tables against other harnesses carry its own caveat: they ran on its suite, and no common standard exists.
What this means
When the verifier is objective, a swarm of agents can be pointed at a measurable direction and left to run, and the cost becomes sandbox-hours and tokens. The environment vendors’ product is the part this rewrite did not need to buy: a verifier for work where parity cannot be compiled.
FAQ
What is Prime Agent?
Prime Intellect’s open coding agent harness, launched in August 2026, which runs as a long-lived daemon with a worker process per session. Prime reports more than 300,000 downloads and over 8 trillion tokens processed. The Rust version adds Windows support in beta and isolates sessions.
Why did the hillclimb have no numeric targets?
Because, in the authors’ words, a fixed threshold tends to become a stopping point. The loop was given a direction, faster and smaller, and a benchmark suite with noise checks, and it kept merging improvements for three days rather than halting at a number.
How were the parity checks run?
Each task’s output went to a separate verifier agent that compiled it and ran parity checks in a fresh sandbox, not the working agent’s own. A dedicated TUI verifier compared terminal frames from the TypeScript and Rust builds, and benchmark runs on fresh 4-core, 8 GB sandboxes withheld results that failed noise checks.