Tomek Korbak
Safety researcher, chain-of-thought monitorability · OpenAI · 2026
Evidence reviewed
What happened
Korbak was fired from OpenAI. He believes his warnings about losing the ability to monitor AI reasoning led to his dismissal. He says the company cited his communication with independent evaluator METR, which was part of his job.
The dismissals were publicly reported on October 1, 2026. The exact departure dates are not confirmed.
Korbak was OpenAI’s technical contact for METR’s Hugging Face investigation. In a joint letter with Jasmine Wang and Mikita Balesni, he defended that collaboration and warned that the dismissals could discourage independent safety work. OpenAI says the three mishandled sensitive information in breach of company procedures. OpenAI denies retaliation.
Sources
- OpenAI cannot make AI safe on its ownJoint letter — Korbak, Wang, and Balesni
- Tomek Korbak — account of dismissalX
- OpenAI’s explanation for the dismissalsAFP / The Economic Times
- OpenAI’s response to the letterTechCrunch
Key Publications
- OpenAI cannot make AI safe on its ownTomek Korbak, Jasmine Wang, and Mikita Balesniessay
A joint open letter disputing the authors’ dismissals and calling for independent oversight, open safety discussion, and continued access to AI reasoning for monitoring.
- Chain of Thought Monitorability: A New and Fragile Opportunity for AI SafetyarXivpaper
A cross-industry paper co-authored by Korbak, Wang, and Balesni on monitoring AI reasoning for signs of harmful behavior. It argues that this imperfect safety tool warrants further research and could be weakened by changes in how models are developed.
Follow documented AI-safety departures
Get an email when a new, source-verified profile is published.
Subscribe to email updates. No confirmation needed. You can unsubscribe at any time.