Definition
A prompt-injection attack where the attacker plants malicious instructions for an AI agent to find and follow, but never sees the agent's response or output directly — they only benefit from whatever action the agent is tricked into taking. Honeypot research has caught real attackers running this technique against exposed AI infrastructure in the wild.
Why it matters
It shows prompt injection has moved from a lab curiosity to an active, opportunistic criminal technique being used at scale against internet-facing AI systems.