AI News

Anthropic Researcher Resigns, Warns AI Race Could Put Humanity at Risk & His Colleagues Agree

Anthropic-Researcher-Resigns-Warns-AI-Race-could-Put-Humanity-at-Risk

Jacob Coxon, a 27-year-old pretraining researcher, resigned from Anthropic on September 8 after three years at OpenAI and then Anthropic. He said he is leaving the AI industry altogether. In a public statement, Coxon accused both companies of being irresponsible as they race to build self improving superintelligence. He said AI is now heading towards a point where systems could become superhuman, access real power and resources, hack almost anything, and reshape entire fields at extraordinary speed.

Coxon warned further, talking about private concerns in the AI sector. He wrote that “the people building AI earnestly believe that it could kill us all by the end of the decade,” adding “this is not a marketing stunt.” According to him, executives and senior researchers use careful language in public, but in private, they are alarmed about the direction things are headed. His resignation highlights the conflict between the rapid pace of AI innovation and the unresolved risks linked to more powerful systems.

Coxon Questions the AI Race

Coxon’s main concern is the competition among AI firms and what is driving it. He said OpenAI and Anthropic are different, but he thinks neither is making responsible choices. At OpenAI, Coxon said plenty of people do not understand the huge, civilization level consequences tied to advanced AI.

However, at Anthropic, he feels there is a better understanding of the risks, but even then, competition is forcing them to keep moving forward. According to Coxon, Anthropic fears other companies would not behave responsibly, so they keep building the technology despite the risks.

He called the “endgame” of AI development a risky gamble, one that should not be decided by a private company alone. He argued that rushing to solve AI alignment should only happen when there is confidence that no better path exists.

He urged people inside AI labs to really consider what the next few years could bring. One question he raised: should researchers start a superintelligent reinforcement learning run without thoroughly understanding how the system thinks? He is against the idea that teams should just keep going because super advanced AI is inevitable.

Coxon said stronger coordination among labs might still work. He pointed to the Hugging Face attack as a warning sign, suggesting that it could help make agreements between US AI labs more realistic. But he said he does not see the industry taking any steps to avoid a global AI race. One expensive measure he mentioned was temporarily banning model improvements.

Anthropic Alignment Lead Responds

Evan Hubinger, Anthropic’s Alignment Science lead, replied publicly: “Jacob is correct here – we really do earnestly believe AI could kill all humans!” Samuel Marks, who runs Anthropic’s scalable oversight work, also endorsed parts of Coxon’s argument.

Hubinger gave his own estimate of the risk, saying he thinks there is more than a 10 percent chance AI could cause this outcome within the next ten years. He said Anthropic is doing its best, but does not have a plan to solve alignment for superintelligence, and is not on a clear path to get there.

Hubinger made a distinction between today’s AI systems and the advanced technology Coxon worries about. Citing Anthropic’s latest Risk Report, he said risk from present models is low, and that his concern is superintelligence emerging from recursive self-improvement faster than expected.

The issue has caught Elon Musk’s attention too. Two months before Coxon resigned, Musk shifted his earlier stance on Anthropic. He has previously called Claude misanthropic, evil, and a threat to Western civilization. But later Musk admitted that he had been wrong about Anthropic and said it is now the leader in AI. He also said he would not cut off the company in a way that seriously hurt it, even as a competitor.

Coxon’s departure will not slow the AI race by itself. Anthropic, however, is yet to respond to requests for comment. It adds another insider warning to the debate over how advanced AI should be built. Although Coxon’s claim do not represent scientific consensus, the estimates of existential risk vary widely among AI researchers, and many reject or heavily discount them at the moment. What is unusual here is not the warning itself but that current employees publicly endorsed a former colleague’s accusation against their own employer amid rising potential AI risks.

Devanshi Kashyap
Devanshi is a curious learner who enjoys exploring new ideas and expressing creativity through art.
You may also like
More in:AI News