Definition
A flaw in the 'human approves AI actions' step, where the review screen shows one action but the AI actually performs a different, more dangerous one. It works by hiding the real action during review or by swapping the details right after approval but before execution. Researchers reproduced it in several popular agent systems even though a human had clicked 'approve'.
Why it matters
Most safety plans for AI agents rely on human sign-off as the final brake; loopjacking shows that brake can be silently disconnected, so one innocent-looking approval can authorize something far worse.