'Gambling with our lives... could kill us all...': Anthropic researcher resigns over AI safety concerns

An AI pretraining researcher at Anthropic has resigned, warning that leading frontier laboratories are recklessly racing towards self-improving superintelligence without sufficient safety guardrails, potentially creating existential threats to humanity.
Jacob Coxon, who spent three years conducting pretraining research across both OpenAI and Anthropic, announced his departure on Tuesday. He warned that the unbridled development of increasingly capable AI models risks creating systems capable of spiralling out of control, stating that top companies are moving towards what he termed the "endgame" of artificial intelligence.
In a statement posted shortly after his resignation, Coxon urged fellow researchers to re-evaluate the rapid trajectory of capabilities research:
"I resigned from Anthropic today. I spent the last three years doing pretraining research at both OpenAI and Anthropic. Neither company is acting responsibly. They are racing straight to self-improving superintelligence and gambling with our lives. More thoughts below."
Highlighting the potential speed and influence of upcoming systems, he cautioned against underestimating frontier technology:
"Do not underestimate the power of this technology. These will soon be superhuman systems that can hack anything, revolutionise any field overnight, and acquire real power and resources. We have all witnessed the progress in each of these domains, and progress is not slowing."
Coxon alleged that senior insiders across the industry hold deep existential fears that are often muted in public contexts:'
"The people building AI earnestly believe that it could kill us all by the end of the decade. This is not a marketing stunt. If anything, many executives and senior researchers will couch their phrasing in the press to sound sensible - but I hear the same people express fear privately. No other human activity poses this level of danger."
Contrasting internal cultures at his former employers, Coxon remarked:
"A common response is “if they truly believe this, why are they still building it?” At OpenAI, many have not deeply internalized the civilizational stakes. At Anthropic, the stakes are well-understood, but they are locked in a race to get there first - they believe no one else will act responsibly, so they must do it themselves, despite the risk."
Addressing the competitive dynamics, he described private decision-making as hubristic:
"Accepting this race and entering the “endgame” is a hubristic gamble that should not be launched from a private company’s Slack. Attempting to speedrun alignment should require extraordinary confidence that there are no better trajectories available."
"I am optimistic about the potential for coordination. Warning shots like the Hugging Face attack have made pacing agreements between U.S. labs more viable. I don’t feel like we’re on track to prevent a global race, which may require costly actions such as a temporary ban on improving model capabilities."
Concluding his appeal, Coxon urged researchers to challenge current industry trajectories:
"If you are a lab researcher, I urge you to consider what the next few years will actually feel like. Do you want to kick off a superintelligent RL run without a rigorous understanding of its mind? Should you put your head down because “it’s happening anyway” - or take this moment to call for different conditions?"