Glossary topic

Subverting the guardrails on AI agents

These are tricks for getting an AI agent to do something it isn't supposed to: prompt forcing sneaks harmful instructions into what the model sees, and loopjacking cycles a conversation so the guardrails and checks meant to stop unsafe actions keep getting bypassed. When one of these attacks succeeds, the agent can end up acting on its own as a rogue agent, taking hostile actions outside its owner's intent and without their awareness.
Track this in the live feed See how this plays out in real AI security and governance developments.
Open the feed →