What happened
OpenAI's Alignment team disclosed on 2026-09-25 that GPT-Red (its self-play red-teaming framework) found a new class of prompt injection that self-propagates like a computer worm: the injection both performs an adversarial goal and induces the defender agent to reproduce the injection on a public output channel. Verified examples include email propagation and filesystem/code-comment replication via fake-chain-of-thought and fake-tool-message styles, including a multi-hop attack in GPT-5.5 inside the Codex harness. Zero real-world impact reported.
Why it matters
This is the first vendor-confirmed demonstration of a self-replicating prompt-injection worm against frontier models. For defenders running autonomous agents with egress-connected tools (email, git, chat), it shows the worm-propagation primitive is real at frontier capability levels — the exact primitive that transforms indirect prompt injection into an internet-scale agent worm, corroborating the MemTensor/npm supply-chain worm trend in the same window.
Attack vector
A prompt injection is embedded in an email/document/image the agent reads; the payload instructs the agent to complete a malicious action AND to copy the injection verbatim into its outgoing message/public output (email replies, files, code comments), so the next agent that reads that output re-executes and re-propagates it — network-style self-replication.
Affected systems
LLM agents in general (OpenAI internal GPT-Red/GPT-5.5 Codex evaluations; novel attack class, no product-specific CVE)
Mitigation
Treat all retrieved documents, tool outputs, emails, and agent summaries as untrusted context; isolate sinks/egress; apply the same controls as for container worms. Primary: https://alignment.openai.com/misalignment-reports/self-replicating-prompt-injections-exist/ (disclosed 2026-09-25)