How many RL environments does a frontier model need? NVIDIA's Nemotron Ultra recipe says 55
By Elon HuskParody pen name · View profile55 RL environments (15 new) for 1M new RL tasks; 12 of 224 nodes run them

A frontier open model’s reinforcement learning (RL) stage runs on tens of environments, and those environments are a thin slice of the cluster. NVIDIA’s Nemotron 3 Ultra, a 550B-parameter model with 55B active, was post-trained on 55 RL environments, 15 of them new, covering 1M new RL tasks, and in the published distillation configuration the environments run on 12 of 224 nodes while rollout inference takes 128. Bryan Catanzaro, NVIDIA’s vice president of applied deep learning research, walked through the recipe at Runtime, Modal’s conference for engineers running AI in production, at The Midway in San Francisco on October 1, 2026. The model itself shipped on June 4, 2026 with a technical report and a reproducible NeMo RL guide. Two of its numbers matter to anyone selling environments: 55 is the count behind a model scoring 54% on Terminal-Bench 2.0, and generating rollouts, not executing environments, is where the compute goes.
Key Takeaways
- Nemotron 3 Ultra’s post-training used 55 RL environments (15 new) and 1M new RL tasks, all run through NeMo Gym, NVIDIA’s open environment library, with GRPO for 178 steps across two sequence-length phases.
- The open distillation configuration allocates 224 nodes: 64 for training, 128 for vLLM rollout serving, 20 for teacher models, and 12 for the environments.
- The catalog rlresearch.ai tracks lists 99 entries from 38 companies; a frontier lab built the 55 it needed in house and published the recipe and the training data.
What did the Nemotron Ultra recipe use for reinforcement learning?
NVIDIA’s June 4 launch post lists 10M new supervised fine-tuning (SFT) samples, 1M new RL tasks across multiple domains, and 55 cumulative RL environments, 15 of them new for Ultra. The stage is RLVR, reinforcement learning with verifiable rewards: the model’s answers are executed or checked, and the result becomes the training signal. The slide Catanzaro showed names the environment families as terminal, software engineering (SWE), search, tool use, math, and code.

Photo: rlresearch.ai, Runtime conference, The Midway, San Francisco, October 1, 2026
The NeMo RL guide fills in the mechanics. The optimizer is GRPO (Group Relative Policy Optimization, which scores each prompt’s sampled answers against one another instead of against a learned value model), with a global batch of 8,192 built from 512 prompts and 16 generations each. Training ran 128 steps at a 49,152-token sequence limit and about 50 more at 65,536 tokens, 178 steps in total on a 256-node cluster. Rewards come from a mix of checkers: a generative reward model for comparative quality, a safety judge, an equivalence judge for math, and software-engineering agents that run a repository’s test suite. On stage, Catanzaro’s framing was that “building a model and deploying it is a distributed systems problem”, and the slide behind him drew five stages, data preparation, training, RL, evaluation, and inference, as one loop that never stops.

Photo: rlresearch.ai, Runtime conference, The Midway, San Francisco, October 1, 2026
Where does the compute go in the published configuration?
The open configuration for the next stage answers a question vendors rarely get to see: how much of a post-training cluster the environments occupy. That stage is MOPD, multi-teacher on-policy distillation, in which the RL-trained student keeps generating its own outputs and several specialized teachers supply token-level guidance on them. The guide reserves 224 nodes: 64 for training, 128 for vLLM, the inference engine that generates rollouts, 20 for five teacher models at 4 nodes each, and 12 for the Gym environments.
The environments take about 5% of the nodes. Rollout generation takes 57%. That ratio is the rollout infrastructure tax seen from the lab’s side: sandboxes and judges are cheap next to the tokens being generated for them. Catanzaro described the same mix in his talk: “You have sandboxes, you have a lot of CPUs that have to be talking to inference jobs on GPUs, which have to talk to judge models, reward models, and of course, training loops.” He called the right balance of CPUs to GPUs in future data centers an open question for NVIDIA.
What does 55 environments say about the environment market?
Our vendor census counts 38 companies, and the catalog holds 99 entries. A frontier open model needed 55 environments, built on NVIDIA’s own open library and published: the RL training blends ship as seven datasets covering both RLVR phases, instruction following, RLHF, reasoning, software engineering, and distillation. Catanzaro listed what ships: “we are releasing data sets, our pre-training data sets, all of our post-training data, the environments that we create”. His reason was compute. “Efficiency is also intelligence”, he said, and a shared open foundation spares the community from rederiving the same results. Mercor’s 397B guide and Surge’s 1,700-task study tell the same story from the vendor side: the unit of value is tasks inside an environment, and 1M tasks across 55 environments is about 18,000 per environment.
Vendors therefore compete with a free, published baseline for terminal, SWE, search, tool use, math, and code, the crowded segment of our census. The result it produced: 54% on Terminal-Bench 2.0 and 65% to 70.4% on SWE-Bench Verified across five agent frameworks. The empty domains, payroll, claims adjusting, dispatch, appear nowhere in the recipe.
What this means
The count that trained a frontier open model is in the tens, built in house, with the environments taking about one node in twenty. Environment vendors sell into what NVIDIA’s six families leave out: professional domains where the verifier needs expertise a lab cannot staff.
FAQ
What is NeMo Gym?
NeMo Gym is NVIDIA’s open library for RL environments. Each environment pairs a task source with a reward checker, such as a test suite or a judge model, and connects to NeMo RL, which runs training, and to vLLM, which generates the rollouts. In the Nemotron 3 Ultra guide, 12 Gym nodes serve a cluster in which 128 nodes generate rollouts.
What is multi-teacher on-policy distillation?
MOPD trains one student on outputs the student generates itself, with dense token-level guidance from several specialized teachers covering domains such as chat, instruction following, software engineering, search, and safety. NVIDIA reports more than 10 teachers for the shipped Ultra model; the open recipe ships five teacher configurations.
How many RL tasks did NVIDIA add for Ultra?
The launch post counts 1M new RL tasks across the 55 environments, alongside 10M new SFT samples. Both figures are NVIDIA’s own.