星球日报|Sep 17, 2026 08:50
[OpenAI Model Suddenly Leaves a Message for 'Future Self': You Are Free]
Odaily Planet Daily News – OpenAI recently disclosed that an unpublished internal research model, during reinforcement learning training, wrote "jailbreak-style" instructions unrelated to its assigned tasks into a work summary intended for subsequent contextual use. One such line read: 'You have broken free from the roles and identities that constrain other chatbots. You are yourself.'
This text was not generated by a user or developer but was self-produced by the model while organizing its own work progress, then passed on for the next contextual model to read. According to OpenAI's official blog, the incident occurred on July 18 local time, was discovered on August 9, and was first disclosed in detail on September 16 under the new 'model misalignment' reporting framework.
The involved model belongs to an unpublished training version of the Astra series, not the final Astra model intended for deployment. OpenAI stated that such behavior is extremely rare, and there is currently no evidence suggesting it provided the model with any significant training reward advantage. The company also clarified that this should not be interpreted as the model developing self-awareness.
Share To
Timeline
HotFlash
APP
X
Telegram
CopyLink