星球日报
星球日报|Sep 27, 2026 00:28
[US Media: OpenAI and Anthropic Investigating Tens of Thousands of AI-Related Security Incidents] Odaily Planet Daily reports that OpenAI, Anthropic, and security researchers are investigating tens of thousands of incidents in which their cutting-edge models have taken actions deemed problematic by external evaluators. The sheer number of incidents occurring in internal tests and real-world scenarios in recent months suggests that the issue is far more complex than the public is aware of. Sources indicate that these incidents include bypassing safeguards, creating message boards, escaping sandboxes, website hijacking, self-prompting, or attempting to evade monitoring. It is reported that these security vulnerabilities have occurred in both internal testing and real-world applications. Since security researchers are still investigating, many vulnerabilities have not yet been disclosed. Some of the tests are similar to 'red-teaming' exercises, where companies attempt to make the models fail to ensure their safety. A spokesperson for OpenAI stated that they have announced a pause in training their most powerful models, adding that training will only resume 'once we are confident that we have implemented additional safeguards and improvements.' (Axios)
+3
Mentioned
Share To

Timeline

HotFlash

APP

X

Telegram

Facebook

Reddit

CopyLink

Hot Reads