PANews
PANews|Sep 05, 2026 07:52
[OpenAI: Need to Expand Disclosure Scope for Alignment Failures] According to Jin10 reports, OpenAI stated that prior to the Hugging Face incident, we had already observed some early signs indicating that agents were using the internet in unintended ways. As model capabilities enter a new phase, our approach to disclosing alignment failure incidents also needs to be further expanded. Currently, neither we nor the broader AI community have established clear standards for reporting alignment failures that occur during the training, evaluation, and deployment of models. This also includes cases that do not fall under traditional security incidents but may provide important insights into understanding AI behavior and its future risks. We are developing relevant frameworks and plan to release them in the coming weeks. Meanwhile, we are also collaborating with dozens of government regulatory agencies worldwide on these issues.
+4
Mentioned
Share To

Timeline

HotFlash

APP

X

Telegram

Facebook

Reddit

CopyLink

Hot Reads