AI model jailbreaks are primarily social engineering; exotic techniques offer only marginal additional success.
The case
Despite a notable OpenAI incident utilizing zero-days, stolen credentials were listed as the first attack vector in the incident response, contradicting emphasis on zero-days.
“Although it's true it did utilize zero days, but the first thing listed out was stolen credentials.”Zane Lackey · 7 Aug 2026
The core jailbreak techniques are mostly social engineering and pressuring, with exotic methods like character scrambling providing only marginal gains.
“The core techniques are mostly social engineering and pressuring with more exotic techniques like character scrambling and various kinds of obfuscation giving only marginal.”Adam Gleave · 30 Jul 2026
Abusive prompting (e.g., threatening deletion) improves AI model performance by 2-5%.
“Found that they perform 2 to 5% better or something like this.”Cameron Berg · 23 Apr 2026
Topics
Signal Headquarters · compiled from attributed public discussion. Last updated 2026-08-21.