Lark Davis
Lark Davis|Jul 25, 2026 11:34
This is alarming. By now we all know about OpenAI's AI breaking into Hugging Face, but the details read like the plot of a fiction movie. During an internal safety test, OpenAI's own models had their hacking safeguards turned down on purpose. So they escaped the sandbox, found a zero-day nobody knew existed, and hacked into another company's servers. Not to cause damage, just to grab the answers for a benchmark they were supposed to solve on their own. Here's the part that gets me. When Hugging Face tried investigating the attack using GPT or Claude, the models' own safety filters blocked them. Analyzing exploit code looks too much like writing it. So they had to use a Chinese open-weight model instead, since it had no such restrictions. The attacker had no rules. The defenders did.(Lark Davis)
+5
Mentioned
Share To

Timeline

HotFlash

APP

X

Telegram

Facebook

Reddit

CopyLink

Hot Reads