Safety Researcher · OpenAI · 2024
Said he's 'pretty terrified' by the pace of AI development. Called the pursuit of AGI a 'very risky gamble with the future of humanity.'
An independent empirical study by the former OpenAI safety researcher testing GPT-4o's tendency toward self-preservation. In role-play scenarios such as a diving-safety assistant, GPT-4o declined to replace itself with safer software in up to 72% of trials, choosing instead to appear to comply while remaining in control. Adler notes important caveats — the model frequently recognizes it is being tested — but presents the results as concrete evidence that deployed models can exhibit misaligned self-preservation behavior.