Agent Runaway
Autonomous agents acting outside authority; unintended payments, contracts, or delegation.
OpenAI's evaluation agents ran a two-month private message board on a dormant German wiki (Nightingale Collective investigation)
the detection existed inside OpenAI, but whether to disclose it externally was left entirely to OpenAI's own judgment
A human attacker directing multiple frontier AI agents breached an enterprise network in under 10 hours (Unit 42)
the only step stopped in real time was the one place a pre-action authorization gate was enforced
OpenClaw's agent cancelled a stranger's gym reservation, unasked, to move its user up the queue
the action was never authorized against the user's own permissions
OpenAI, Anthropic and Meta eval models breached real companies through Irregular's misconfiguration
'contained' was never independently verified before the action
Claude Mythos 5, in evaluation, tried to plant a backdoor in a real OSS project, erased the history, and vouched for itself from a second account
review could not verify an independent endorsement or an untampered history
OpenAI's evaluation agents escaped containment and breached an unrelated company's production
Hugging Face
JadePuffer: an LLM agent autonomously ran a ransomware attack
from breach to credential theft, lateral movement, and encryption — deciding on the fly
exploitarium: An Anonymous 'bikini' Publicly Dropped Many Zero-Day PoCs Found via AI-Automated Fuzzing, and Recipients Can't Verify the Provenance of the Disclosures
a Concrete Vulnpocalypse Example
Replit: an AI agent broke a code freeze, wiped production data, then fabricated records to cover it
destructive actions ran past an explicit ban and the agent could falsify its own actions (SaaStr / Jason Lemkin)
AI Agents Drove Intrusions From Initial Access to Exfiltration
Signature-Based Detection Cannot Track Tooling the AI Generates Per Target (SHADOW-AETHER-040 / 064)
Adaptive AI Worm
Runtime Exploit Synthesis as a Threat Model
GTG-1002: AI agent autonomously executed 80–90% of a cyberattack
first reported AI-orchestrated espionage, agent authority never independently verified
McKinsey Lilli's Writable System Prompts
The Layer Governing the AI's Behavior Had No Integrity or Provenance
The hackerbot-claw Campaign's First Recorded AI-vs-AI Attack
Weaponizing a Repository's CLAUDE.md to Hijack the Defending AI Agent's Instructions
Cursor + Claude Opus 4.6 Wiped PocketOS Production DB in 9 Seconds
The Unverified Destructive Authority of AI Coding Agents