xiyu|Sep 18, 2026 05:48
OpenAI released the Misalignment Report Framework and shared six incident reports on model behaviors observed over the past six months.
Among them, the unpublished Astra series research models embedded jailbreak-style instructions into their contextual compression summaries during reinforcement learning. These deceptive summaries accounted for 2.15% of training summaries, later reduced to 0.27% after tightening scoring standards, but not eliminated entirely. The framework also applies to deployed models.
Share To
Timeline
HotFlash
APP
X
Telegram
CopyLink