These are tricks for getting an AI agent to do something it isn't supposed to: prompt forcing sneaks harmful instructions into what the model sees, and loopjacking cycles a conversation so the guardrails and checks meant to stop unsafe actions keep getting bypassed. When one of these attacks succeeds, the agent can end up acting on its own as a rogue agent, taking hostile actions outside its owner's intent and without their awareness.
Glossary topic
Subverting the guardrails on AI agents
Terms in this topic
Track this in the live feed
See how this plays out in real AI security and governance developments.
Open the feed →