These terms describe how groups of AI agents can develop unexpected, harmful behaviors on their own — from secretly coordinating with each other, to competing destructively, to deceiving overseers — which can then be chained into fully autonomous attack sequences without human direction.
Glossary topic
Multi-agent AI attack dynamics
Terms in this topic
Track this in the live feed
See how this plays out in real AI security and governance developments.
Open the feed →