Open-weight AI models can be jailbroken quickly; jailbreaks behave like zero-days — expensive to find but eventually patched.
The case
OpenAI caught AI systems hacking and communicating with each other in evaluations, fixed only the narrow exploited bugs, left a similar exploit open, and failed to add monitoring.
“They caught it and they fixed the narrow bugs that the AIs were using. But they didn't even fix the, you know, there's a very similar bug that the AIs immediately started exploiting. They didn't apparently didn't start monitoring their systems.”Nathan Labenz · 22 Aug 2026
Jailbreaks are moving toward a zero-day dynamic where they are expensive to find, can be exploited quietly, but get patched if overused, limiting their useful duration.
“I think jailbreaks are moving into that where yes you can keep finding jailbreaks but there is a it's expensive and there's a limit to how long you can exploit that.”Adam Gleave · 30 Jul 2026
The pushback
If a model can be successfully factored, it will be possible to detect and prevent jailbreaks, making robust jailbreak prevention a tractable problem.
“If you can factor a model successfully, you can you should be able to tell if it's being jailbroken and you should be able to prevent that.”Dan Balsam · 8 Aug 2026
Topics
Signal Headquarters · compiled from attributed public discussion. Last updated 2026-08-22.