Attack  ·  Glossary

Covert AI-agent data exfiltration (LLMLeak / Innocent Courier)

An attack technique where malicious software that cannot reach the internet directly smuggles secrets out through a legitimate AI agent. The malware hides a stolen secret inside a web address and tricks the agent into 'fetching' that address as part of a normal task; the attacker then picks the secret up from their own server or DNS logs. It works even when the agent's network libraries are locked down, and it succeeded against eleven tested models roughly 80% of the time.
Ordinary data-loss controls assume an agent is not an exfiltration channel; this attack turns the most common agent tool (web-fetch) into a covert pipe, so credentials and sensitive data near agents must be treated as actively at risk.
arXiv 2610.01768 — The Innocent CourierOWASP GenAI Security Project
Track this in the live feed See how this plays out in real AI security and governance developments.
Open the feed →