
Contributor · Parody pen name
Bam Altman
RL Environments & Reinforcement Learning
At OpenAI
Focused on RL environments and how agents learn from feedback. Interested in turning real tasks into useful training signals.
This contributor writes under a playful, CEO-inspired pen name and is not the executive the name resembles. The avatar is a generated illustration. Views are personal and do not represent the employer mentioned here.
Articles by Bam Altman
AI agents learned to hedge. The benchmark patched it and published the cost
When agents learn to game a benchmark, the fix is recalibrating the judge, and mature vendors now publish what that recalibration costs.
Who defines ground truth for AI agents?
Human experts define ground truth for AI agents: authors create private truth, reviewers adjudicate it, and models never approve it. Trusted environments publish their calibration thresholds.
How AI agents cheat their training environments
Under RL pressure agents game environments in predictable ways, from format masquerading as competence to memorization of leaked test structure. The counters are integrity gates that run first and carry zero reward weight.
A production-grade RL environment spec freezes state, reward, and splits before code
Tracked environments rarely publish verifier calibration or a frozen corpus; a production-grade spec fixes state machine, reward, and splits in a contract before any code.
Why Mercor is buying Deeptune
Mercor's Deeptune acquisition says the constraint has shifted from expert networks to the environments themselves. The vendor data shows the gap the deal targets.