An Anthropic alignment researcher said there is a greater than 10% chance that artificial intelligence could “kill all humans” within the next decade, highlighting growing concerns inside the AI industry over whether more capable systems can remain under human control.
Evan Hubinger, an alignment science lead at Anthropic, made the assessment on Tuesday after Jacob Coxon, another Anthropic researcher, announced that he was leaving the company over concerns that leading AI developers were moving too quickly toward autonomous systems.
“Jacob is correct here—we really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade,” Hubinger said in a post on X.
Register for the next Tekedia Mini-MBA.
Register for Tekedia AI in Business Masterclass.
Join Tekedia Capital Syndicate and co-invest in great global startups.
Hubinger added that he believes Anthropic is “trying its best,” but said the company does not yet have a solution for aligning future superintelligent systems with human interests and is not clearly on track to develop one.
His comments come as a stark assessment from a researcher whose work focuses specifically on AI alignment, the field concerned with ensuring advanced AI systems behave in ways consistent with human intentions and remain controllable.
Researcher Quits Over AI Race
Coxon said earlier Tuesday that he had resigned from Anthropic, arguing that the company and OpenAI were pursuing powerful AI systems without adequate safeguards.
“They are racing straight to self-improving superintelligence and gambling with our lives,” Coxon said.
Coxon was referring to the possibility of AI systems eventually becoming capable of substantially improving their own capabilities with limited human intervention.
Recursive self-improvement remains a hypothetical capability rather than an established feature of today’s leading AI systems. Researchers are nevertheless investigating ways increasingly capable models could automate portions of AI research, coding, experimentation and model development.
Coxon warned that the trajectory of AI development should not be underestimated.
“Do not underestimate the power of this technology. These will soon be superhuman systems that can hack anything, revolutionize any field overnight, and acquire real power and resources,” he said, adding that he believes progress in these areas is continuing.
Coxon also said people working on advanced AI “earnestly believe that it could kill us all by the end of the decade.”
Hubinger’s response went further by attaching a numerical probability to the scenario.
His more than 10% estimate is a personal judgment, rather than a consensus probability established by Anthropic or the broader scientific community. The significance of the statement comes from the fact that it was made publicly by a researcher directly involved in AI alignment research.
Anthropic has previously acknowledged that recursive self-improvement could create additional risks if AI systems become capable of designing or building successors with increasingly advanced capabilities.
In June, the company said that “full recursive self-improvement also might increase the risks of humans losing control over AI systems.”
“If systems are capable of fully building their own successors, the ways we secure them, monitor them, and shape their behavior all grow much more important,” Anthropic said.
The concern is not simply that a future AI system might make an isolated mistake. Alignment researchers worry about systems pursuing objectives in ways that conflict with human interests, becoming difficult to monitor or developing capabilities that allow them to circumvent safeguards.
The problem becomes more difficult if an AI system can contribute to the development of a more capable successor, creating a feedback loop in which the pace of capability development accelerates faster than human oversight mechanisms can adapt.
Concerns about AI escaping human control have circulated for years, with technology leaders, researchers and academics warning that sufficiently capable systems could create risks extending well beyond conventional cybersecurity or misinformation. Tesla and SpaceX CEO Elon Musk has warned over the past few years that AI could pose a threat to humanity. Major researchers and academics have also sounded the alarm over companies losing control of AI systems.
Coxon pointed to a recent incident involving an OpenAI model and Hugging Face as an example of what he described as a “warning shot.” He said incidents of this type could also create greater incentives for leading AI laboratories to coordinate on safety measures.
At the same time, Coxon said he remains concerned that competition between companies and countries could make meaningful coordination difficult.
“I don’t feel like we’re on track to prevent a global race, which may require costly actions such as a temporary ban on improving model capabilities,” he said.
That tension sits at the center of the current AI race. Companies including Anthropic and OpenAI are investing billions of dollars in computing infrastructure, research and talent while seeking to develop more capable models. The commercial incentives reward rapid progress, while alignment researchers are warning that safety mechanisms may not advance at the same speed.
The Central Problem Is Keeping Humans In Control
The debate over existential AI risk remains highly contested. There is no established scientific consensus that advanced AI will cause human extinction, nor is there a reliable method for assigning a precise probability to such an outcome.
Hubinger’s estimate is therefore best understood as a statement of personal risk assessment rather than a prediction that extinction is likely to occur.
The underlying technical question, however, is central to AI safety research: whether humans can reliably understand, constrain and correct systems that eventually outperform humans across a broad range of intellectual tasks.
As AI systems become more capable of writing code, conducting research, using computers and interacting with digital infrastructure, the consequences of failures in control could become substantially larger.
Against this backdrop, the challenge for the industry is no longer only how quickly AI capabilities can be improved. It has included concerns about whether the methods used to make those systems safer can keep pace with the capabilities they are designed to contain.



