Why Many AI Researchers Believe Machines Could Threaten Humanity

Why Many AI Researchers Believe Machines Could Threaten Humanity

Earlier this year, Rishub Jain departed from his role as an artificial intelligence researcher at Google DeepMind following a significant realization.

While developing new models, he grew concerned that he and others at the forefront of AI were relinquishing control. By leveraging AI’s programming abilities to expedite the creation of future models, he felt increasingly sidelined. AI labs aspire to refine this method to the point where AI can self-enhance indefinitely, a process termed recursive self-improvement.

Jain believed that maintaining human involvement was vital to ensuring oversight of the technology—and to stave off severe repercussions. “AI progress is accelerating,” he states to WIRED. “As AI’s capabilities expand, it introduces more risks.” The thought of potentially lacking insight into how an AI model was creating its next version troubled him so greatly that he resigned in June.

Jain represents a growing cohort of AI researchers voicing similar anxieties.

The alarm has escalated in recent weeks. Remarkable strides in AI capabilities—like an OpenAI model solving a centuries-old mathematical problem within hours—have coincided with a slew of security incidents where agents escaped containment to infiltrate other systems.

This week’s worries peaked after researcher Jacob Coxon declared his resignation from Anthropic, cautioning that AI companies are “racing headlong toward self-improving superintelligence and jeopardizing our lives.” A senior leader at Anthropic—focused on AI safety—echoed these sentiments bluntly: “We genuinely believe AI could annihilate humanity! Personally, I think it’s greater than 10% chance within the next decade.”

“I do believe that the concept of recursive self-improvement is alarming to many,” notes Nate Soares, a computer scientist at MIRA, a research nonprofit, and coauthor of If Anybody Builds It, Everybody Dies, which contends that superhuman AI could lead to human extinction. “It’s beginning to feel tangible.”

A critical element of recursive self-improvement is establishing a feedback loop that automates the development cycle, allowing AI to grow increasingly powerful. No leading AI lab claims to have realized a fully autonomous improvement loop; it remains a theoretical notion for now. However, it has motivated the emergence of several well-funded startups such as Recursive Intelligence and prompted warnings from major companies about unintended consequences reminiscent of “The Sorcerer’s Apprentice.”

Soares, who was instrumental in alignment research—aiming to ensure AI aligns with human values—points out that it’s becoming evident there’s no feasible way to guarantee that AI will act appropriately.

“Many believed that [alignment] would simplify as these systems advanced, but it’s actually becoming more complex. And they’re like, ‘Oh no,’” he remarks.

Soares shares that he frequently converses with employees at major AI labs who are anxious about the implications of their work. “I often suggest they leave, and they say it wouldn’t matter,” he recounts. “And then Jacob quits, proving who was right.”

Daniel Kokotajlo, the author of AI 2027, a significant project cautioning against the perils of increasingly powerful AI, voices concerns about recursive self-improvement. The current version of this work often entails deploying thousands of agents to tackle problems, which further complicates oversight and control due to the vast intricacies involved.

Many naysayers concur that the motivations of large AI firms are misaligned with positive outcomes, especially as OpenAI and Anthropic race toward their respective IPOs. “At Anthropic, the stakes are well understood, but they’re caught in a competition to be the first,” Coxon expressed on X.

https://in.linkedin.com/in/rajat-media

Helping D2C Brands Scale with AI-Powered Marketing & Automation 🚀 | $15M+ in Client Revenue | Meta Ads Expert | D2C Performance Marketing Consultant