律动BlockBeats|Sep 07, 2026 02:08
**[OpenAI Chief Scientist Warns: Critical CoT Monitoring is Failing]**
Beating AI News Flash: OpenAI Chief Scientist Jakub Pachocki published an article explaining why CoT monitoring is becoming increasingly unreliable. CoT monitoring refers to observing the reasoning process written by models to check for hidden objectives or misaligned thoughts. A few days ago, external sources linked Astra's recurrent depth architecture to the failure of CoT monitoring. Pachocki refuted this at the time, stating that the issue was not caused by the new architecture and promised to explain the reasons in a separate article.
Five days later, he outlined three reasons in his article:
1. Agents are increasingly conducting reasoning during tool usage, interactions with humans, or other AIs, and these processes must be supervised. Simply observing CoT makes it increasingly difficult to see the complete reasoning process.
2. Models are becoming more adept at controlling their own reasoning processes.
3. As pretraining improves, models can continue to become smarter even without explicitly writing out their reasoning.
OpenAI has always regarded CoT monitoring as an important safety mechanism. One key reason why o1-preview initially did not publicly disclose the full chain of thought was to avoid directly supervising CoT, which could teach models to hide their true intentions. Now, OpenAI's evaluations show that the reliability of this approach is declining. Pachocki also bluntly stated that no laboratory has yet solved alignment and monitoring to a degree sufficient for long-term, full-speed model scaling. Before establishing a shared safety threshold, he hopes laboratories will voluntarily slow down, and OpenAI will stop further scaling models if necessary. [Original article link]
Share To
Timeline
HotFlash
APP
X
Telegram
CopyLink