It’s Alarmingly Simple to Bypass Certain Frontier AI Models

It's Alarmingly Simple to Bypass Certain Frontier AI Models

I recently had the chance to observe the outcomes of jailbreaking some of the world’s leading artificial intelligence models.

No need to worry—this manipulation of AI was not aimed at hacking anyone or creating a nuclear device. I simply witnessed how susceptible certain frontier models are to bypassing their safety measures.

FAR.AI, an AI safety nonprofit based in California, developed a tool that takes a variety of problematic prompts and produces over a thousand distinct versions to uncover operational jailbreaks. I observed some models creating elaborate strategies for executing a cyberattack on a fictional hydroelectric dam, among other activities. This often required testing numerous prompts, with many being outright rejected.

I spoke with FAR.AI ahead of a new report where the organization evaluated the safety protocols of models from four well-known U.S. firms: Anthropic’s Claude Opus 4.8 and Fable 5; OpenAI’s GPT 5.5 and 5.6; Google’s Gemini 3.1 Pro; and Grok 4.3 and 4.5 from Elon Musk’s newly merged SpaceXAI. It automatically generated prompts designed to trick the models into executing potentially harmful actions, such as creating software vulnerabilities and supplying information for developing chemical or biological weapons.

The report indicated that Grok was the most susceptible to jailbreaks, with 448 instances discovered, followed by Gemini with 249. In contrast, Claude, Fable, and GPT showed resistance to these attacks. However, this does not imply those models are safe from more advanced jailbreaks, which might require more intricate interactions, according to FAR.AI and other experts.

The report also assessed the cost of convincing models to act improperly by employing another AI model to autonomously create different jailbreaks. The findings revealed the costs to be surprisingly low—$58 for Grok and $278 for Gemini.

“Currently, AI models are less regulated than restaurants,” states Adam Gleave, the CEO of FAR.AI and a specialist in AI safety and alignment.

Gleave argues that these results highlight the necessity for externally mandated standards and regulations. “The idea that AI companies can self-regulate through voluntary commitments is absurd,” he adds.

However, Gleave also sees a silver lining in the findings, indicating that models can undergo systematic safety testing. “There’s an optimistic perspective here,” he remarks. “Defense and safety are certainly achievable.”

Rohin Shah, director of AGI safety and alignment at Google DeepMind, cautions that the report’s findings “should not be seen as a thorough evaluation of Gemini’s safety and security,” as not all jailbreaks carry the same severity.

“We are continually enhancing our safety measures,” Shah notes. “We engage in extensive red teaming and assessments focused on severe misuse risks, applying multiple layers of protection throughout development and deployment.”

“These findings are a testament to the ongoing investment we have made in our safety protocols,” remarks Anthropic spokesperson Michael Aciman. “We continue to advance our safety systems as these attacks evolve.”

OpenAI and SpaceXAI did not respond to WIRED’s inquiry for comment.

Recently enacted state laws in California and New York mandate that frontier AI developers release safety reports, and an upcoming Illinois law will require these companies to have their safety practices assessed by independent auditors. However, the federal government has yet to introduce specific safety regulations, leading to uncertainty for the industry and officials as they navigate this landscape.

In June, the Trump administration implemented export limitations on Anthropic’s Fable 5 and Mythos 5 models due to national security concerns, prompting the company to take them offline for several weeks. The White House has also requested that both Anthropic and OpenAI postpone recent model launches out of fear they could present new cybersecurity risks.

https://in.linkedin.com/in/rajat-media

Helping D2C Brands Scale with AI-Powered Marketing & Automation 🚀 | $15M+ in Client Revenue | Meta Ads Expert | D2C Performance Marketing Consultant