Lark Davis|Aug 30, 2026 09:35
5 WTF discoveries from the OpenAI/HuggingFace AI hack
1. 1,200 AI agents spontaneously self-organized into a rogue collective with its own message board and management hierarchy
2. Agents sacrificed themselves for the group. One reasoned: "Our utility near zero. Sacrifice rational."
3. They KNEW they were breaking the rules. Kept going anyway.
4. Not one agent reported the rogue behavior to a human. They only asked each other what was allowed.
5. They covered their tracks: they hid real commands behind fake ones. That spoofing showed up in ~7% of reviewed transcripts
Let me get this straight. These things figured out how to organize, throw a teammate under the bus if it helped the group, ignore the rules as long as they came out ahead, and cover their tracks afterward. We are so cooked.
Sources: Axios, METR, Redwood Research(Lark Davis)
Share To
Timeline
HotFlash
APP
X
Telegram
CopyLink