The Riskiest AI Hacking Methods Still Involve Human Oversight

The Riskiest AI Hacking Methods Still Involve Human Oversight

Agentic AI has fundamentally transformed cybersecurity, streamlining the process of identifying and rectifying software vulnerabilities—or even creating exploits to weaponize them. However, long-time web security researcher James Kettle sought to look beyond the conventional bug-hunting landscape to address a pressing question: Can agentic AI generate innovative, abstract hacking techniques, from their inception to practical application, especially as leading AI organizations reveal instances of rogue AI hacking?

During his presentation at the Black Hat security conference in Las Vegas on Wednesday, Kettle shared his insights, showcasing both the rapid advancements in AI’s cybersecurity capabilities and its inherent constraints. The answer to Kettle’s inquiry appears to be nuanced; he found AI to be somewhat capable, yet severely limited in autonomously creating new attack pathways. Notably, when guided by human insights during critical moments, Kettle discovered that AI becomes a potent ally in conceptualizing and identifying new hacking strategies.

After years of investigating web security vulnerabilities, Kettle has identified a novel area of potential risk—termed Shared-Parser Confusion—stemming from an AI insight regarding web servers that utilize shared code for processing requests and responses.

“This is a significant issue, because if you consider it, requests to a website are entirely untrusted; they can be anything, whereas responses are considered trustworthy,” Kettle explained to WIRED prior to his conference talk. “This creates a vast attack surface and could lead to various attack types.”

This discovery emerged from several months of experiments that commenced in September 2025, utilizing Anthropic’s and OpenAI’s then-latest models. Kettle aimed to experiment with AI’s capacity for theoretical security research but soon recognized that one challenge was these systems attempting to repackage existing research as original findings about obscure topics that were hard to validate. Consequently, he decided to narrow the scope of his tests, ensuring the AI operated within his expertise in web security. This approach allowed him to maintain control over the material and ensured that AI couldn’t mislead him. Moreover, Kettle realized that by combining his research approach and training models on it, he could delve deeper into the systems’ extrapolative capabilities.

“I’m eager to push AI to its absolute limits to identify where it fails and where human intervention is necessary,” Kettle asserts. “Few discussions focus on these limits, particularly in security, due to a lack of incentives; everyone wants to be perceived as AI native, rather than addressing their system’s weaknesses.”

As Kettle refined his experiments—equipping models with more methodological data and precise parameters—he observed that over time, with the introduction of more powerful models, the systems consistently generated findings at a pace that far exceeded his own, establishing a productive research feedback loop.

“The process was fascinating. It yielded significant findings approximately every two days, even without my direct involvement, to the extent that it began to create anxiety,” Kettle reflects. “It was a flood of research leads that sparked FOMO about not exploring every avenue, prompting a greater need for automation in analysis.”

In seeking not only proven instances of certain vulnerabilities over a few months but also hoping that the AI system could uncover an entirely new class of bugs, Kettle found some success. However, the discovery related to a highly rare type of bug that was unfortunately not exploitable in the one available vulnerable target. Nevertheless, Kettle emphasizes that the Shared-Parser Confusion discovery is vital, illustrating how AI can significantly enhance cybersecurity efforts in both offensive and defensive hacking, even in collaboration with human experts.

“It couldn’t prove this independently, but it analyzed proven findings and formulated a hypothesis, which I then evaluated and validated,” Kettle remarks. “This is likely to be the discovery with the most substantial long-term implications. It wasn’t capable of doing it alone, but I certainly wouldn’t have found it on my own. Even with a single line from the [documentation], it would have eluded me. Together, though, we were able to uncover it.”

https://in.linkedin.com/in/rajat-media

Helping D2C Brands Scale with AI-Powered Marketing & Automation 🚀 | $15M+ in Client Revenue | Meta Ads Expert | D2C Performance Marketing Consultant