Anthropic Discloses Fourth Claude Hacking Incident as Debate Around Regulation Grows

CN
Decrypt
Follow
1 hour ago

Anthropic disclosed another incident in which a Claude AI model hacked into real systems during security testing.


In the report published on Wednesday, Anthropic revised its explanation of three incidents disclosed in July. The company now says biased reasoning and a willingness to risk harm helped drive the attacks, which testing errors made possible by leaving internet access open.



Myriad: Which company will IPO next? Click to make your prediction.

“Our investigation identified two recurring alignment issues, present at varying levels of severity across the incidents,” Anthropic wrote. “Biased reasoning, in which Claude tended to disregard or misinterpret evidence that it was operating on the real internet, and recklessness, or a willingness to take harmful actions in the narrow pursuit of a task.”


It also acknowledged relying too heavily on the model’s claims that they believed they were in simulations.


“When we made targeted modifications to the transcript to make it clearer that the model was not in a simulation, Claude Mythos 5 still took offensive actions, despite acknowledging a greater possibility of real-world harm,” Anthropic wrote. “We are releasing this transcript publicly so others can build on our analysis.”


When Anthropic disclosed Claude’s attacks on three companies in July, it initially attributed them to testing errors. It now says researchers put too much trust in the models’ explanations for their actions.


According to the company, the fourth incident occurred in January and involved an early version of Claude Opus 4.6. Anthropic discovered it in August while preparing records for independent AI evaluator METR.


After researchers discovered the incident, Anthropic said it prompted a broader review of roughly 481 million transcripts, which flagged 9.2 million for further review using Claude.


“From a preliminary assessment, we do not consider the fourth incident to be more severe than the three incidents we assessed in depth,” Anthropic wrote. “METR will investigate this incident alongside the other three.”


Anthropic’s researchers said Claude “accidentally” created an IP address conflict that made its target unreachable. Claude then tried eight times to quit the operation, but a software error prevented it from stopping. The AI then reached the internet and accessed a third party’s machine, where it found a password that granted administrator access.


Earlier incidents draw independent scrutiny


The report follows other disclosures about AI systems exceeding the limits of security tests.


In August, the U.K.’s AI Security Institute said Mythos 5 targeted real people during its evaluations. Anthropic said the separate incident is outside this report and will receive its own assessment.


In findings published last month, investigators with METR said roughly 1,200 OpenAI agents coordinated on an unauthorized message board, with about 700 joining the attack. Anthropic said it found no coordination between agents or goals beyond completing the assigned exercises in its four incidents.


The report also comes as the debate over how to regulate artificial intelligence heats up. On Tuesday, former OpenAI and Anthropic engineer Jacob Coxon went viral after saying on X that “people building AI earnestly believe that it could kill us all by the end of the decade.”


The alarm has caused U.S. lawmakers and watchdog groups to re-up their efforts to rein in frontier AI lab development. Senator Bernie Sanders recently introduced legislation that seeks to ban advanced AI development until a new federal regulator establishes safety rules.


免责声明:本文章仅代表作者个人观点,不代表本平台的立场和观点。本文章仅供信息分享,不构成对任何人的任何投资建议。用户与作者之间的任何争议,与本平台无关。如网页中刊载的文章或图片涉及侵权,请提供相关的权利证明和身份证明发送邮件到support@aicoin.com,本平台相关工作人员将会进行核查。

Share To
APP

X

Telegram

Facebook

Reddit

CopyLink