
Contributor · Parody pen name
Satya Nutella
RL & Post-Training
Previously at Microsoft Research
Focused on reinforcement learning and supervised fine-tuning. Interested in how post-training changes model behavior.
This contributor writes under a playful, CEO-inspired pen name and is not the executive the name resembles. The avatar is a generated illustration. Views are personal and do not represent the employer mentioned here.
Articles by Satya Nutella
Anthropic's AI agents beat 28 human experts at alignment research, and humans became the baseline
Automation has reached the loop's last human stage: expert researchers now serve as the measured baseline, and the human role narrows to defining failures and judging what counts as better.
1,700 RL tasks lifted Kimi K2.7 on every coding benchmark tested
Broad expert task portfolios, not benchmark targeting, are now the reliable training recipe: one 1,700-task set lifted five external coding benchmarks at once and made the agent faster as well as stronger.
RL environments just moved a 397B model. The training guide is public
Expert-built RL environments moved a 397B open-weight model by 70 percent relative on held-out professional tasks, and the recipe is published rather than proprietary.
Training AI on office work made it better at coding
Environment training transfers across domains. A model post-trained on office tasks with no coding in the mix gained 5.8 points on SWE-Bench Pro, which changes what labs are buying.