Member of Technical Staff (Manager), Scalable Oversight · Anthropic · 2026
Evidence reviewed
Manager of Anthropic's Scalable Oversight team who left in August 2026 to join the independent evaluator METR. Benton said frontier AI companies were racing toward recursively self-improving systems while underinvesting in safety, and warned that progress could become uncontrollable. He called for slower development, disclosure of safety incidents and capability gains, minimum safety standards, and independent assessments.
Benton wrote on September 11, 2026 that he had left Anthropic's safety team two weeks earlier. This establishes an August 2026 departure, but the exact day is not publicly documented.
“Competition pushes every frontier company to underinvest in safety; the cost of falling behind is too high.”
Benton framed his concern as industry-wide. He said safety researchers inside frontier companies face a competitive trap: slowing down may cede ground to less cautious rivals, while continuing may contribute to severe harm. His move to METR was intended to strengthen independent evaluation and public visibility into frontier risks.
Benton and his co-authors test whether reasoning models faithfully disclose the cues that influence their answers. They find that chain-of-thought monitoring can reveal some undesired behavior but frequently misses the true influence of prompt hints, making it insufficient to rule out rare, catastrophic behavior.
Benton and his co-authors develop threat models and evaluations for whether advanced systems could subvert capability testing, behavioral monitoring, or deployment decisions. Their tests of Claude 3 Opus and Claude 3.5 Sonnet suggest minimal mitigations were then sufficient, while stronger and more realistic safeguards would likely be needed as capabilities improved.
0 open forecasts · 1 warning and contemporaneous claims
AI agents smarter than any human could emerge within the next couple of years
“Within the next couple of years, we may be sharing the world with AI agents smarter than any human alive today.”
Get an email when a new, source-verified profile is published.
Confirm your address by email. You can unsubscribe at any time, and we never share your address with third parties.
Benton explains why he left Anthropic's safety team for independent evaluation work at METR. He argues that competition pushes frontier companies to underinvest in safety, warns that recursively self-improving systems could make progress uncontrollable, and calls for transparency, incident reporting, minimum safety standards, and independent assessment.