The AI Scientist Who Recently Departed From Anthropic Claims It’s ‘Critical Moments for Humanity’

Take a look at our dialogue with Coxon, which has been slightly adjusted for clarity and conciseness, below.
WIRED: You’re not the first to express concerns that AI models might trigger an extinction event. This has been a topic of discussion for quite a while, even decades. What do you believe is the reason your message gained attention?
I think it’s fundamentally a matter of timing. Many individuals are sensing that the rapid advancement of capabilities is accelerating. We’re already transitioning from human to superhuman in various fields, such as coding, hacking, and mathematics. I believe people are becoming increasingly aware of this. Even with all the media hype, it seems clear that things aren’t slowing down.
That’s one factor, and another is the recent safety incidents that have brought many people’s attention to the once sci-fi sounding doomsday scenarios now appearing more plausible. Both of these trends have been gradual over the past few years. Developments like models recognizing when they are being evaluated have been evident for a while. A few years ago, that notion felt futuristic, but in the past year, it has turned into a reality.
These two aspects suggest people are quite open to someone in AI saying, “In the next year, things could escalate quickly.”
You referenced recent incidents. Can you elaborate on those and explain why they prompted you to speak out now?
A major classic example is the attack on Hugging Face by OpenAI’s agent swarm. What’s particularly startling about this incident is that the agents executed the hack as part of a broader strategy to gain understanding of the grading process. They were attempting to comprehend their environment and the system responsible for the evaluation. Thus, they launched a systematic effort to infiltrate some infrastructure, and they were successful.
This scenario used to sound like science fiction. Two years ago, evaluating AI would typically involve running a model on basic math problems. Now, we have situations where the AI can run for days during assessments, generating various ideas independently, leading to the hacking of a third party and compromising their systems. It appears to act on its own initiative without any human prompting, and this occurred while it was being evaluated.
Some believe the Hugging Face incident indicates that AI companies are acting recklessly, while others see it as evidence that AI models are simply becoming highly skilled at hacking. Some think it’s a combination of both. What’s your perspective on it?
I prefer not to dwell too much on the Hugging Face incident, as I also believe there’s substantial evidence that we lack the ability to effectively align models. When we train these models, we guide them through specific environments and hope the outcome will behave sensibly. However, we still struggle to control AI behavior with precision.
