infrastructure 5 min read

Prime Intellect ran 30 million microVM sandboxes in beta. RL sandboxes are now a priced product

About 30M sandboxes created in beta; 20,000+ concurrent; $0.02 per vCPU-hour plus $0.0125 per GiB-hour

About 30 million microVM sandboxes created in beta, with more than 20,000 running at once

Hardware-virtualized sandboxes for agentic reinforcement learning (RL) are now a metered product. Prime Intellect launched Prime Sandboxes on September 23, 2026: microVMs, each with its own guest Linux kernel, that its researchers and early customers had created about 30 million times during a closed beta, with a dashboard showing more than 20,000 running at once. Pricing is per resource, $0.02 per vCPU-hour, $0.0125 per GiB of memory per hour, and $0.0002 per GiB of disk per hour, which the company calls a third of what other large sandbox providers charge, through December 22. The design matches the pattern our reference spec adopted in August: a hardware virtualization boundary rather than a container, tunnels for inference traffic, and direct integration with the training loop. The two features that spec listed as unsolved, state snapshots for per-episode reset and sandbox forking, are on Prime’s roadmap rather than in the release.

Key Takeaways

  • Shipped: a microVM per sandbox with its own guest kernel, Docker and background jobs inside, tens of thousands of concurrent sandboxes, integrations with Prime’s verifiers library, prime-rl trainer, and Hosted Training, and Prime Tunnels to reach inference without network setup.
  • Priced: $0.02 per vCPU-hour, $0.0125 per GiB-hour of memory, $0.0002 per GiB-hour of disk, no subscription or minimum, promotional through December 22. A 2-vCPU, 4 GiB, 10 GiB sandbox costs about $0.09 an hour.
  • Roadmap: GPU microVMs, state snapshotting, sandbox forking, and shared persistent workspaces. Per-episode database reset depends on the second and third.

What did Prime Intellect ship?

The unit is a microVM rather than a container. “Each sandbox is a fully capable Linux virtual machine”, in Prime’s words, with its own guest kernel, so Docker, background services, and kernel-dependent workloads run inside it unchanged. Prime’s comparison table sets that against gVisor, the user-space kernel approach, which it says has documented syscall and subsystem gaps and needs complex setup to run Docker at all. The isolation classes are the ones a June 2026 comparative security study separated on attack surface and leakage: hardware virtualization, user-space kernel, and plain container.

Prime Sandboxes dashboard for a 24-hour window: 20,292 concurrent sandboxes, 865,133 created, $38,302 estimated spend, 0.0% error rate, with a concurrency chart holding near 18,000 to 24,000
The dashboard Prime published with the launch: 20,292 sandboxes live, 865,133 created in the previous 24 hours, and $38,302 of estimated spend for that day, about 4.4 cents per sandbox created. The beta total it reports is about 30 million.

Chart: Prime Intellect, September 23, 2026 · source

The dashboard screenshot is the most concrete number in the launch: 865,133 sandboxes created in one 24-hour window for an estimated $38,302, which works out to about 4.4 cents per sandbox, with concurrency holding between roughly 18,000 and 24,000 and an error rate of 0.0%.

The RL integration is the distinguishing claim. Sandboxes plug into Prime’s verifiers library, its prime-rl trainer, and its Hosted Training service, with access to more than 365,000 prebuilt environments on its hub, and teams can bring their own Docker image without a new packaging format. Prime Tunnels give a sandbox a path to inference services without network configuration. The company says it built scheduling around RL’s burst pattern, “caching images near the compute and preferentially scheduling sandboxes” so that thousands start within seconds.

What does an episode cost?

Prime Sandboxes at launch: six capabilities shipped, four on the roadmap, with per-episode reset depending on the roadmap items

At list prices a sandbox with 2 vCPUs, 4 GiB of memory, and 10 GiB of disk costs $0.04 plus $0.05 plus $0.002, about $0.09 an hour. Terminal-Bench 4.0 set every task to an 8-hour agent timeout on August 28, so a maximal episode at that size costs about $0.74 in sandbox time, before the inference tokens that dominate the bill. That ordering is the rollout infrastructure tax quantified by Graviet et al.: the execution substrate is a real cost, and a small one next to generation. Pricing it per vCPU-hour with no minimum makes it a line item an environment vendor can pass through, and the launch asks “3x cheaper than other providers” to be read as a market with incumbents.

Which open problems remain open?

Our August spec named two. The first is per-episode application-database reset: a deterministic verifier needs every episode to start from the same seeded state, and resetting the VM is not enough when an HR system or a general ledger lives inside it. Prime lists state snapshotting and sandbox forking as near-future work, which is the mechanism that problem needs, and marks neither as shipped. The second is an action-level event schema for desktop sessions; the release is Linux-only, GPU microVMs are also roadmap, and Windows-native work, where much of the enterprise workload lives, is not mentioned.

Prime’s own framing of why fidelity matters is the right one for RL: “Silent differences from production can be more dangerous than hard failures because they can reward behaviors that do not transfer”. A container that silently lacks a subsystem teaches the agent a workaround, and the reward signal cannot tell the workaround from competence. That is the training-side version of the goal misgeneralization problem: a correct reward over a subtly wrong environment still trains the wrong policy.

What this means

The isolation layer of an RL environment is now something a vendor rents by the vCPU-hour rather than builds. What stays in the environment vendor’s hands is the part above it: the seeded application state, the per-episode reset, and the verifier, which is where the margin was already.

FAQ

What is a microVM?

A lightweight virtual machine with a minimal device model that boots in well under a second while keeping a hardware-virtualization boundary between the workload and the host. Each Prime sandbox runs its own guest Linux kernel, which is what lets Docker and kernel-dependent software run inside it.

Why not containers or gVisor?

Containers share the host kernel, and gVisor interposes a user-space kernel that Prime says leaves documented syscall and subsystem gaps. For an agent that runs arbitrary commands for hours, a hardware boundary is the appropriate isolation class, and full kernel fidelity avoids the silent environment differences that can reward behavior which does not transfer.

Does Prime Sandboxes support Windows or GPUs?

Not at launch. GPU microVMs are listed as near-future work alongside state snapshotting, sandbox forking, and shared persistent workspaces. Windows is not mentioned in the release.