Definition
A jailbreak technique where an attacker hides malicious instructions inside an encrypted or encoded payload so that an AI model's safety filters — which typically scan plain text — never see the harmful content before decrypting and acting on it. It has been shown to defeat both the input and output safety checks of major chatbots with no user clicks required.
Why it matters
It proves that today's safety guardrails can be systematically bypassed with a repeatable trick rather than a one-off bug, meaning any AI product using similar filtering needs re-testing against this class of attack.