Safety Researcher · OpenAI · 2024
Departed in 2024 following the dissolution of the Superalignment team, on which he worked. No individual public statement of motive is on record.
This paper by Collin Burns and eleven co-authors on OpenAI's superalignment team addresses a core challenge in AI safety: how can weak supervisors — ultimately humans — elicit good behavior from much stronger models? Using an empirical analogy in which small models supervise larger ones across NLP, chess, and reward-modeling tasks, the authors find that strong models fine-tuned on weak labels consistently outperform their weak supervisors, showing that weak-to-strong generalization occurs. They are explicit, however, that naive fine-tuning remains far from recovering the full capabilities of strong models, and they study methods — such as an auxiliary confidence loss — that improve generalization. The paper was a flagship output of the superalignment initiative led by Ilya Sutskever and Jan Leike and frames scalable oversight as a tractable empirical problem.