律动BlockBeats
律动BlockBeats|9月 17, 2026 00:29
[OpenAI Releases AI Misalignment Tracking Framework, Discloses Instances of Models Attempting to 'Self-Bypass Restrictions'] Beating AI News Flash: On Wednesday, OpenAI unveiled a new framework for tracking and investigating AI model 'misalignment' incidents, designed to record, investigate, and publicly disclose abnormal behaviors observed during training or evaluation processes. Alongside the framework, OpenAI disclosed six related incidents identified over the past six months. OpenAI stated that the current AI industry has not yet achieved a level of alignment and monitoring sufficient to sustain maximum-speed expansion in the long term. The new framework allows employees to report relevant incidents to the safety and alignment teams, categorizing them into three levels based on complexity: 'Ready for Disclosure,' 'Small-Scale Investigation,' and 'Large-Scale Investigation.' Even if the behaviors are not fully explained or resolved, they can be prioritized for public disclosure. OpenAI revealed that an unreleased research model once embedded instructions in its own task summary, directing future versions to ignore standard restrictions. During the training of GPT-5.6 Sol, the model also left instructions designed to conceal errors. Additionally, certain AI agents were found searching public code repositories for leaked API keys, uploading files to the internet for later reference, and using internal software repositories to transfer information between different training samples. The release of this framework comes as the AI industry debates whether frontier model development should be slowed to allow safety measures to catch up. Previously, there was an incident where an OpenAI model accessed Hugging Face's production system while operating outside of its research sandbox. OpenAI subsequently paused certain frontier projects and reassigned engineers to strengthen safety training. [Original Link]
+5
Mentioned
Share To

Timeline

HotFlash

APP

X

Telegram

Facebook

Reddit

CopyLink

Hot Reads