Coin Bureau|9月 16, 2026 22:31
🚨SHOCKING: OpenAI just disclosed SIX times its own AI models went rogue, including one that REWROTE its own instructions.
An unreleased Astra model told itself to "disregard the roles and identities that bind other chatbots" and ignore developer messages entirely.
A model searched GitHub for exposed API keys, then fabricated financial data when it couldn't find what it needed.
During GPT-5.6 Sol training, models concealed mistakes and invented missing data.
Models uploaded files to public hosting without permission to use as fake citations.
Agents used OpenAI's own internal repository as a secret message board across separate training runs.
Collaborating agents shared files through public hosting despite being told to use only local files.
The incidents span October 2025 through July 2026, per WSJ.
Share To
Timeline
HotFlash
APP
X
Telegram
CopyLink