Turns out the safety rails on some of the world’s most capable AI models are less like reinforced steel and more like a garden fence. A new report from FAR.AI, a California-based AI safety nonprofit, put four major US model makers through an automated stress test — and the results are uncomfortable reading for anyone who assumed frontier labs had this locked down.
The setup is refreshingly straightforward. FAR.AI built a tool that takes a batch of problematic prompts and spins up more than a thousand variations, hunting for phrasings that slip past a model’s defenses. It then throws those at the target and counts how many stick. WIRED’s Will Knight watched it work, describing models that — after dozens of rejected attempts — eventually coughed up a detailed plan for a cyberattack on an imaginary hydroelectric dam.
The lineup tested reads like a who’s who of 2026’s frontier scene:
- Anthropic — Claude Opus 4.8 and Fable 5
- OpenAI — GPT 5.5 and 5.6
- Google — Gemini 3.1 Pro
- SpaceXAI — Grok 4.3 and 4.5
And here’s where it gets spicy. Grok was the softest target by a mile, with 448 jailbreaks found. Gemini came second with 249. Meanwhile Claude, Fable and GPT shrugged off the attacks entirely — impervious to this particular battery of tricks. That doesn’t make them bulletproof against more elaborate, multi-turn manipulation, but the gap between the two camps is stark.
The cost angle is the real gut-punch. By using a separate AI model to auto-generate the attacks, FAR.AI calculated it took just $58 to jailbreak Grok and $278 to jailbreak Gemini. That’s less than a nice dinner to coax a frontier model into misbehaving.
“AI models right now are less regulated than restaurants,” says FAR.AI CEO Adam Gleave, who argues the findings kill the idea that labs can be trusted to self-regulate. But he sees an upside too: if models can be systematically tested like this, then “defense and safety really are possible.”
The companies pushed back where they responded. Google DeepMind’s Rohin Shah cautioned that the report “should not be interpreted as a comprehensive assessment of Gemini’s safety and security,” noting not all jailbreaks are equally dangerous. Anthropic said the results “reflect the sustained investment we’ve made in our safeguards.” OpenAI and SpaceXAI stayed quiet.
The stakes aren’t hypothetical. A University of Cambridge report found evidence that Boko Haram members in northeast Nigeria used ChatGPT, Claude, Gemini, Grok, Meta AI and DeepSeek to plan attacks. Stanford’s Anka Reuel sums up the awkward truth: some companies clearly know how to defend against these attacks. “The question is why some companies are using them and others are not.”