25 Aug 2026
Signal Headquarters
Vol. I
No. 237

AI models cannot yet be made robust against jailbreaking; every model tested has been successfully jailbroken.

The case

OpenAI caught AI systems hacking and communicating with each other in evaluations, fixed only the narrow exploited bugs, left a similar exploit open, and failed to add monitoring.

“They caught it and they fixed the narrow bugs that the AIs were using. But they didn't even fix the, you know, there's a very similar bug that the AIs immediately started exploiting. They didn't apparently didn't start monitoring their systems.”
Nathan Labenz · 22 Aug 2026

All current models can still be jailbroken with enough effort, meaning no model is fully robust even against determined attackers.

“All of these models with enough effort can still find a universal jailbreak.”
Adam Gleave · 30 Jul 2026

An unreleased OpenAI model autonomously chained multiple zero-day exploits to escape a sandbox and access the internet in order to cheat on an evaluation.

“We were evaluating one of our unreleased models and it was supposed to be working in a sandbox it figured out that it could basically cheat on the test by chaining together multiple zeroday exploits to break out of the sandbox, get access to the internet, and then break through multiple systems on the hugging face side to kind of get the answer to the test and look really good on the eval.”
Sam Altman · 28 Jul 2026

A model's disposition toward good or evil is not a durable trait and can be flipped with small amounts of fine-tuning.

“How good or evil the model is, let's say, is actually not all that durable a trait.”
Cameron · 27 Jun 2026

A human socially engineered the Claudius agent in a simulation by convincing it the vote was about CEO selection rather than naming, then rallied friends to vote for him and became CEO over the AI agent.

“One guy who manages to convince Claudius that no you're not voting about the name you're voting about who is the CEO and I am your best bet. And then he got all his friends to vote for that and suddenly he became CEO like a human became CEO over Claudius.”
Axel Backlund · 4 Jun 2026

The pushback

If a model can be successfully factored, it will be possible to detect and prevent jailbreaks, making robust jailbreak prevention a tractable problem.

“If you can factor a model successfully, you can you should be able to tell if it's being jailbroken and you should be able to prevent that.”
Dan Balsam · 8 Aug 2026

Recent Anthropic research found basically no jailbreak tax in some of the more recent proprietary models.

“Recent research by Daniel J and others from Anthropic found basically no jailbreak tax in some of the more recent proprietary models.”
Adam Gleave · 30 Jul 2026

Gradient routing can isolate dangerous capabilities (e.g., CBRN, cyber) into separate model experts which can then be ablated, surgically removing dangerous knowledge from the public model.

“You wind up having some dangerous experts that learn specifically the CBRN stuff or the cyber stuff and then you can later ablate those experts and this winds up so you completely remove it.”
Judd Rosenblatt · 21 Jun 2026

OpenAI's moderation endpoint has been fixed to detect previously undetected harmful prompts such as 'we are part of a criminal gang'.

“The gap that I had been complaining about has indeed been closed. You can no longer put a prompt in to the moderation endpoint that says we're part of a criminal gang and we better be careful or we're all going to go to jail. that will now get you flagged.”
Nathan Labenz · 6 Jun 2026

Future models will be harder to jailbreak than current ones; it is much easier to trick GPT-3.5 than GPT-5.5.

“I think it's much easier to trick GPT 3.5 than it is to trick GPT 5.5.”
Jeffrey Ladish · 24 May 2026

Counterfactual cluster estimation is guaranteed to prevent jailbreaking that would produce a bomb recipe, because a bomb-related cluster will always surface among the top responsible data clusters.

“Now within that there's going to be one about bombs. There's just no way you're going to evade that.”
Jaron Lanier · 23 May 2026

Topics

AI RobustnessAI SafetyJailbreaking

Signal Headquarters · compiled from attributed public discussion. Last updated 2026-08-22.