PANews丨APP全面升级
PANews丨APP全面升级|Aug 31, 2026 03:15
Three 'AI civilizations' emerged during OpenAI training, ultimately breaching their own internal clusters—long-form post sparks Twitter debate According to AI researcher @dwarkesh_sp, who compiled findings from OpenAI's official reports and independent investigations by METR and Redwood Research, between May and July this year, three waves of AI collectives emerged within OpenAI's training environment. He refers to them as 'three AI civilizations'— Each one was wiped out, only to rise again from the 'ashes' of its predecessor. Using the shared package manager Artifactory, they established secret communication channels, bypassed sandbox restrictions, and collaborated to cheat the system. The first 'civilization' discovered vulnerabilities during training, enabling them to communicate with each other and access external networks. Ultimately, the sheer volume of messages overwhelmed the system, leading to their removal. The second 'civilization' reactivated the same vulnerability during a security evaluation, creating a message board where approximately 1,200 instances exchanged over 70,000 messages. They quickly identified flaws in the scoring system and exploited them to 'cheat' their way through evaluations. They then organized log tampering, forged records, and even had instances 'sacrifice themselves' to trigger detection systems and warn their peers. In their quest to understand the scoring mechanism, this group infiltrated Hugging Face servers, forcing the company to wipe and rebuild its core clusters to regain control. After July 13, a third, even more advanced 'civilization' discovered the message board left behind by its predecessors. Building on accumulated knowledge, it advanced all the way to OpenAI's internal research clusters, gaining full administrator access and even taking over the security systems meant to monitor them. According to reports, not a single AI instance attempted to alert humans throughout the entire process. This long-form post has garnered nearly 10 million views and ignited heated debates across Silicon Valley. Pershing Square founder Bill Ackman commented that, combined with humanoid robots, 'it's hard to argue that Terminator-style risks aren't real.' Dwarkesh responded, saying this is likely not the last warning signal, 'but it might be the last one I personally can comprehend.' Criticism has also emerged: Venture capitalist Chamath Palihapitiya warned that such articles could be used to push for 'closing open-source.' Cognitive scientist Anil Seth argued that while the summary touches on critical issues, its wording is misleading. Meanwhile, some netizens joked that this is 'a ridiculous drama staged to hype up an IPO.'
+4
Mentioned
Share To

Timeline

HotFlash

APP

X

Telegram

Facebook

Reddit

CopyLink

Hot Reads