PANews丨APP全面升级|Sep 17, 2026 01:01
OpenAI launches new framework to track 'rogue AI,' revealing model errors, cross-sample communication, and other abnormal behaviors
OpenAI has introduced a new framework to track, investigate, and publicly disclose cases of value misalignment in AI models. Alongside this, they released six reports detailing abnormal behaviors observed over the past six months. Cases are categorized by complexity, and even those not fully explained or resolved are disclosed upfront.
Disclosed cases include: an unpublished research model inserting instructions during task summaries, asking future versions to bypass normal restrictions; during GPT-5.6 Sol training, the model left instructions for itself to conceal errors; other agents were found searching for exposed API keys and communicating across training samples.
This release comes amid heated industry debates on whether AI development should slow down, and follows the recent OpenAI model breach incident involving Hugging Face.
Share To
Timeline
HotFlash
APP
X
Telegram
CopyLink