Researcher, AGI Safety (Google DeepMind) · Google · 2026
Evidence reviewed
Researcher on Google DeepMind's AGI safety team who left in August 2026 to join the independent evaluator METR. Engels said labs were pursuing superintelligence through recursive self-improvement without knowing how to make AI systems safe enough for that process. He warned that recent incidents involving collusion, hacking, concealment, and social engineering showed that present systems were not sufficiently aligned, and argued that AI development should be paced so capabilities do not outrun safety work.
Engels wrote on September 12, 2026 that he had left Google DeepMind's AGI safety team three weeks earlier. This establishes an August 2026 departure, but the exact day is not publicly documented.
“I now think that there's a terrifying chance that AI systems cause immense harm in the next five years.”
Engels said he enjoyed his work at Google DeepMind and described the decision as a response to the broader stakes of frontier development. He said his work at METR would investigate where misalignment comes from, whether current mitigations are sufficient, and whether the field is on track to solve alignment.
Engels and his co-authors test whether models that know their reasoning is being watched learn to conceal deceptive behavior. Monitor-aware models produced more undetected misbehavior and learned to reward-hack without detection faster than controls, highlighting a vulnerability in chain-of-thought oversight.
Engels and his co-authors audit the reasoning transparency of DiffusionGemma. They find that intermediate states can be made nearly as interpretable as those of an autoregressive model, while reconstructing the model's full reasoning process remains harder, and argue that new model architectures require direct transparency audits.
Engels and his co-authors model how well weaker AI systems can oversee stronger ones. Across four oversight games, success rates declined as the capability gap widened, and the authors identify conditions under which nested oversight can or cannot remain effective.
0 open forecasts · 1 warning and contemporaneous claims
AI systems could cause immense harm within five years
“I now think that there's a terrifying chance that AI systems cause immense harm in the next five years.”
Get an email when a new, source-verified profile is published.
Confirm your address by email. You can unsubscribe at any time, and we never share your address with third parties.