Anthropic Researcher Resigns, Warns of >10% AI Extinction Risk

Jacob Coxon resigned from Anthropic, warning labs are racing toward self‑improving superintelligence and “gambling with our lives.” Anthropic’s alignment lead estimates >10% extinction risk within a decade.

Jacob Coxon, a researcher who worked on pretraining at both OpenAI and Anthropic, announced his resignation and said he will leave the AI industry for ethical reasons. On social media he wrote, “They are racing straight to self‑improving superintelligence and gambling with our lives.” He added that under aggressive scenarios development could be “out of control” by the end of next year.

Evan Hubinger, head of Alignment Science at Anthropic, responded on the same platform and called Coxon’s concerns real. Hubinger wrote that he personally places the chance that advanced AI could kill all humans above 10 percent over the next ten years. He framed that number as his own estimate, said he views present models as posing relatively low risk, and identified a future risk from systems that improve their own design.

Researchers describe recursive self‑improvement as a process in which AI tools assist or automate research and engineering tasks to design and tune successor systems. That process could speed capability gains in ways that are hard for humans to predict or control.

Anthropic has told investors and the public that recursive self‑improvement is not inevitable, but the company has also warned it could arrive faster than regulators and institutions are prepared for. Hubinger wrote that Anthropic is “trying its best” on alignment work while acknowledging the company does not yet have a plan to solve alignment for a superintelligence and is not clearly on track to do so.

Coxon argued that teams inside leading labs recognize the stakes but face competitive pressure that rewards moving faster, because slowing down could hand an advantage to rivals. The resignation and Hubinger’s post highlight that tension between safety work inside companies and market incentives to advance capabilities.

Other researchers reacted to Coxon’s account. Gary Marcus described the resignation as “compelling and informed from the inside,” while noting he did not agree with every part of the argument. Anthropic continues to fund alignment research, publish risk reports and implement safeguards, and Coxon’s departure has prompted renewed attention to how companies balance safety research with competitive development.

Articles by this author