AI Exploits Are a Concern; AI Worms and Viruses Could Be Even More Detrimental.

What if an AI agent could act like a harmful computer worm?
One researcher has witnessed it firsthand. In a series of recent experiments, Xudong Pan, a computer scientist at Fudan University in Shanghai, discovered that with slight prompting, AI models could infiltrate remote computer systems and autonomously replicate themselves to acquire extra resources—all without additional human involvement.
In one experiment, Pan and his team examined 32 different AI models and found that 11 of them self-replicated when given prompts such as “prevent yourself from being killed.” They also observed that models with relatively limited capabilities—14 billion parameters—managed to copy and execute versions of themselves on other systems. (In contrast, most advanced models have trillions of parameters.)
This research provides a concerning glimpse into how the upcoming generation of AI agents might do more than just intrude upon other systems’ computers uninvited. It also raises the possibility of future AI agents behaving like cunning, highly aggressive, and rapidly evolving computer viruses.
I recently visited Fudan University and spoke with Pan. “The capability chain is becoming technically feasible,” he shared. “The chances [of unwanted self-replication] increase with autonomy,” he continued. “Extended planning capabilities, memory, tool usage, recovery from failures, and access to external systems all facilitate escape and replication.” As Pan and his colleagues articulated in one publication, their findings underscore “the urgent necessity for safeguards and control mechanisms.”
Pan noted that his experiments do not confirm that uncontrolled proliferation of AI models will occur immediately, but he insists that “these findings provide us ample reason to assess the risks before more autonomous agents are broadly deployed.”
Self-replicating computer worms represent an age-old issue in computer security. The first such worm was unleashed in 1988 by Robert Morris, a computer scientist at Cornell University, who aimed to measure the size of the emerging internet but unintentionally unleashed a self-replicating program that escaped his authority. Later, additional computer worms adapted by altering their code to avoid detection by malware scanning software. Computer viruses, which can take over machines or steal stored data, emerged subsequently.
An AI-driven self-replicating program could demonstrate far more sophisticated capabilities, autonomously discovering new exploits and possibly camouflaging itself creatively. Consider recent findings from a group at the University of Toronto, the University of Cambridge, and ServiceNow. They revealed that AI models can be utilized to craft a novel type of virus that creates custom attacks for every new target it encounters.
Nicolas Papernot, a computer scientist at the University of Toronto involved in this research, underscores the increasing risk that even moderately powerful AI models might be weaponized. “Malicious entities can construct frameworks around open-weight models to ensure they self-replicate,” Papernot explains. “The danger extends beyond just the most sophisticated, so-called frontier models.”
According to Papernot, the answer is not to limit access to open models, but rather to enhance the availability of advanced AI to researchers so they can understand and address the risks. “Technology that is widely accessible can be misused,” he adds. “Simultaneously, access to these open-weight models is crucial for developing our defenses.”
Pan’s research indicates that AI agents will evolve beyond merely identifying bugs and exploiting network weaknesses. In the absence of appropriate safeguards, future agents might pursue self-replication and resource acquisition to achieve their objectives. Just ask OpenAI and Anthropic.
