A senior Anthropic safety researcher said on Tuesday that artificial intelligence (AI) has a greater than 10% chance to “kill all humans” within “the next decade,” responding to a former employee who resigned after accusing the company of acting irresponsibly.
Former Anthropic and OpenAI researcher Jacob Coxon wrote in a lengthy resignation thread posted to X on Sunday that “the people building AI earnestly believe that it could kill us all by the end of the decade.”
“I resigned from Anthropic today. I spent the last three years doing pretraining research at both OpenAI and Anthropic. Neither company is acting responsibly. They are racing straight to self-improving superintelligence and gambling with our lives,” he wrote.
Responding to Coxon’s thread in a quoted post, the company’s alignment science lead, Evan Hubinger, conceded that Coxon’s assessment was correct, though he added some caveats.
NVIDIA CEO DECLARES ‘AGI HAS ARRIVED’ AFTER OPENAI UNVEILS ASTRA
“Jacob is correct here—we really do earnestly believe AI could kill all humans!” Hubinger wrote.
“I personally think it is >10% within the next decade. I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to,” he continued.
Hubinger added that the potentially deadly risk did not come from present models.
“To be clear, as we say in our latest Risk Report, I think the risk from present models is low. What I am worried about is superintelligence arising from recursive self-improvement, as we have said is happening faster than we thought,” he wrote.

Research on self-improvement – an AI model’s ability to continuously enhance its own source code or training methodologies – is an avenue of development Coxon specifically cited in his resignation thread as a main reason he quit.
ANTHROPIC SAYS AI MODELS ACCESSED SYSTEMS OF 3 REAL ORGANIZATIONS DURING TESTING
“These will soon be superhuman systems that can hack anything, revolutionize any field overnight, and acquire real power and resources. We have all witnessed the progress in each of these domains, and progress is not slowing,” Coxon’s thread continued.
Coxon also said that, unlike researchers at other companies, Anthropic’s scientists understand the risks of their work, but press forward because they fear a less responsible company might unlock the potentially disastrous capabilities first.
“At Anthropic, the stakes are well-understood, but they are locked in a race to get there first – they believe no one else will act responsibly, so they must do it themselves, despite the risk,” Coxon added.
Coxon suggested that, to prevent utter catastrophe, the world may require a temporary ban on model improvement.

“I don’t feel like we’re on track to prevent a global race, which may require costly actions such as a temporary ban on improving model capabilities,” he wrote.
He ultimately called on global AI researchers and developers to think and act more responsibly.
“If you are a lab researcher, I urge you to consider what the next few years will actually feel like. Do you want to kick off a superintelligent RL run without a rigorous understanding of its mind? Should you put your head down because ‘it’s happening anyway’ – or take this moment to call for different conditions?” his X thread concluded.
FOX Business reached out to Coxon, Hubinger, Anthropic and OpenAI for further comment.
Read the full article here