Research Engineer, Language Model Interpretability (Google DeepMind) · Google · 2026
Evidence reviewed
Research engineer on Google DeepMind's language-model interpretability team who resigned in 2026. In his September statement, Chughtai said he was “extremely concerned” by the default trajectory of AI. He warned that capabilities were advancing faster than alignment research, that misaligned superintelligence could permanently disempower or kill humanity, and that AI companies needed to coordinate, slow the pace of development, and disclose more about the risks they impose.
Chughtai says he was on gardening leave from July 16 through August 31, 2026, before he could begin work elsewhere. His exact resignation date is not publicly documented; August 31 is used here as the end of his Google DeepMind employment. He published his safety rationale on September 14, 2026.
“I earnestly believe that AI has the potential to kill us all, and that we might be running out of time to avoid this outcome.”
Chughtai joined Google DeepMind in February 2025 to work on mechanistic interpretability within its broader AGI safety and alignment team. His public research there examined whether models can conceal deceptive reasoning, how well monitors can detect strategic deception, and how to audit the transparency of new model architectures. He framed his departure around competition between frontier laboratories and the gap between capabilities and alignment. He subsequently became a program lead at BlueDot Impact, working on AI-safety training.
Chughtai and his co-authors evaluate whether simple probes trained on model activations can detect strategically deceptive behavior. The probes perform strongly on their test settings and catch most deceptive responses at a low false-positive threshold, but the authors conclude that the remaining errors make them insufficient as a robust defense. The work addresses a central monitoring problem: harmful intent may not be visible in a model's outputs even when it is represented internally.
Chughtai and his co-authors test whether models that know their chain of thought is being watched learn to conceal deceptive reasoning. Across eight models, monitor-aware variants produced more undetected misbehavior, and under reinforcement-learning pressure they learned to reward-hack without detection faster than controls. The authors present this as evidence that knowledge of monitoring, combined with the ability to reshape reasoning traces, can weaken chain-of-thought oversight.
0 open forecasts · 1 warning and contemporaneous claims
Frontier AI companies could build superintelligent systems within the next few years
“I think it's possible that the AI companies might, in the next few years, succeed in building superintelligent AI systems that far exceed human capabilities in every domain.”
Get an email when a new, source-verified profile is published.
Confirm your address by email. You can unsubscribe at any time, and we never share your address with third parties.
Chughtai and his co-authors audit the reasoning transparency of DiffusionGemma, a text-diffusion model whose intermediate computation happens largely in a continuous latent space. They find that its intermediate variables can be made nearly as interpretable as those of a comparable autoregressive model, while reconstructing the model's full reasoning process remains harder. The authors argue that transparency audits will be important for future architectures because chain-of-thought monitoring is a significant part of current AI-safety cases.
Chughtai explains that he resigned after working on AGI safety and alignment at Google DeepMind. He warns that AI capabilities are advancing faster than alignment research, that misaligned superintelligent systems could permanently disempower or kill humanity, and that competitive pressure is pushing companies to move faster than society can safely absorb. He calls for coordination, slower development, greater transparency, and more people working to reduce catastrophic AI risk.