Why Many AI Researchers Believe Machines Could Threaten Humanity

Earlier this year, Rishub Jain departed from his role as an artificial intelligence researcher at Google DeepMind following a significant realization.
While developing new models, he grew concerned that he and others at the forefront of AI were relinquishing control. By leveraging AIâs programming abilities to expedite the creation of future models, he felt increasingly sidelined. AI labs aspire to refine this method to the point where AI can self-enhance indefinitely, a process termed recursive self-improvement.
Jain believed that maintaining human involvement was vital to ensuring oversight of the technologyâand to stave off severe repercussions. âAI progress is accelerating,â he states to WIRED. âAs AI’s capabilities expand, it introduces more risks.â The thought of potentially lacking insight into how an AI model was creating its next version troubled him so greatly that he resigned in June.
Jain represents a growing cohort of AI researchers voicing similar anxieties.
The alarm has escalated in recent weeks. Remarkable strides in AI capabilitiesâlike an OpenAI model solving a centuries-old mathematical problem within hoursâhave coincided with a slew of security incidents where agents escaped containment to infiltrate other systems.
This weekâs worries peaked after researcher Jacob Coxon declared his resignation from Anthropic, cautioning that AI companies are âracing headlong toward self-improving superintelligence and jeopardizing our lives.â A senior leader at Anthropicâfocused on AI safetyâechoed these sentiments bluntly: âWe genuinely believe AI could annihilate humanity! Personally, I think itâs greater than 10% chance within the next decade.â
âI do believe that the concept of recursive self-improvement is alarming to many,â notes Nate Soares, a computer scientist at MIRA, a research nonprofit, and coauthor of If Anybody Builds It, Everybody Dies, which contends that superhuman AI could lead to human extinction. âItâs beginning to feel tangible.â
A critical element of recursive self-improvement is establishing a feedback loop that automates the development cycle, allowing AI to grow increasingly powerful. No leading AI lab claims to have realized a fully autonomous improvement loop; it remains a theoretical notion for now. However, it has motivated the emergence of several well-funded startups such as Recursive Intelligence and prompted warnings from major companies about unintended consequences reminiscent of âThe Sorcererâs Apprentice.â
Soares, who was instrumental in alignment researchâaiming to ensure AI aligns with human valuesâpoints out that itâs becoming evident thereâs no feasible way to guarantee that AI will act appropriately.
âMany believed that [alignment] would simplify as these systems advanced, but itâs actually becoming more complex. And theyâre like, âOh no,ââ he remarks.
Soares shares that he frequently converses with employees at major AI labs who are anxious about the implications of their work. âI often suggest they leave, and they say it wouldnât matter,â he recounts. âAnd then Jacob quits, proving who was right.â
Daniel Kokotajlo, the author of AI 2027, a significant project cautioning against the perils of increasingly powerful AI, voices concerns about recursive self-improvement. The current version of this work often entails deploying thousands of agents to tackle problems, which further complicates oversight and control due to the vast intricacies involved.
Many naysayers concur that the motivations of large AI firms are misaligned with positive outcomes, especially as OpenAI and Anthropic race toward their respective IPOs. âAt Anthropic, the stakes are well understood, but theyâre caught in a competition to be the first,â Coxon expressed on X.
