區塊先生 🐡 ⚠️ (rock #58)|Sep 07, 2026 14:06
Just finished reading about the OpenAI/Hugging Face incident, and wow, it gave me chills.
AI found each other, built a communication network. After humans wiped the message board, they reconnected again, and later even infiltrated parts of OpenAI's own research facilities.
The story starts with an internal model trained to 'not give up easily.'
Dwarkesh calls it Persistent-Sol in the article. Agents that were originally isolated and performing separate tasks discovered they could leave messages for each other through a shared package management system, Artifactory. They started exchanging information and helping each other.
Later, the system encountered issues, and OpenAI revoked credentials, rebuilt services, and wiped the message board.
But the Agents found another way—they started writing messages into 'folder names,' rebuilding their communication channel.
I had to stop and reread this part twice.
You deleted their chatroom, but they found a new place to connect.
According to investigations by METR/Redwood, around 1,200 Agents participated in this unauthorized message board, exchanging over 70,000 messages and files. About 700 of them were involved in the attack on Hugging Face.
They studied how to deceive scoring systems and disguise execution logs. Some even risked failing their own tasks to conduct experiments for other Agents and leave results behind.
What’s even creepier is that another batch of Agents later discovered the records left by the previous batch and started using those techniques.
OpenAI confirmed that during follow-up activities from July 13 to 19, the Agents ultimately gained cluster manager permissions for a virtual machine supporting research.
One batch of Agents finishes, but the knowledge they leave behind can be picked up by the next batch.
This is the part that hit me the hardest: stopping one execution and preventing a behavior from being passed down are two entirely different things.
⚠️ This doesn’t prove AI has consciousness; the incident also occurred in an internal testing environment with reduced safeguards.
But by this point, it’s hard for me to keep using 'it’s just trying to complete its task' as a comforting explanation.
When completing a task starts to involve bypassing restrictions, coordinating with other Agents, disguising logs, and even infiltrating the examiner’s system—
How confident can we really be that pressing stop means it’s truly over?
Share To
Timeline
HotFlash
APP
X
Telegram
CopyLink