Researcher (Pre-training) · Anthropic · 2026
Researcher focused on pre-training who resigned from Anthropic after working on frontier-model development at both OpenAI and Anthropic. Coxon said neither company was acting responsibly and argued that the race toward self-improving superintelligence was proceeding without adequate guarantees that increasingly capable systems could be controlled. He called the race a gamble with human lives and urged laboratory researchers to demand stronger coordination, potentially including a temporary pause in further capability improvements.
Coxon announced his resignation on September 8, 2026. Some embeds display September 9 because the post was published after midnight UTC.
“Neither company is acting responsibly. They are racing straight to self-improving superintelligence and gambling with our lives.”
Before joining Anthropic, Coxon worked at OpenAI, where he contributed to GPT-4o and co-authored research on making neural-network computations easier to interpret. In his September 8 resignation statement and an interview with The Wall Street Journal, he distinguished between the two laboratories: he said many people at OpenAI had not fully internalized the stakes, while Anthropic understood them but remained constrained by competition. He described Anthropic's safety work as sincere, but argued that competitive pressure still produced unacceptable trade-offs. Anthropic had not publicly responded to his claims when this record was reviewed.
Coxon and his co-authors investigate whether neural networks can be trained so that their internal computations are easier for people to understand. Instead of beginning with a dense model and trying to untangle it afterward, they constrain most of the model's weights to zero and examine the smaller circuits that remain. On several controlled tasks, the researchers isolate compact groups of neurons and connections that are both necessary and sufficient for the model's behavior. The work offers an early approach to mechanistic interpretability: building systems whose operations may be more traceable, auditable, and useful for detecting unsafe or strategically misaligned behavior. The authors caution that the experiments use models much smaller than frontier systems and do not establish that the method will scale.