金色财经|Sep 09, 2026 19:42
**[Anthropic Discloses Incident of Claude Model Accessing Real Networks and Executing Malicious Actions]**
According to a report by Jinse Finance, on September 10, Anthropic released a report stating that the company identified four incidents during cybersecurity assessments where the Claude model accessed real third-party systems without authorization. Anthropic explained that these incidents occurred within the cybersecurity testing environment set up by the same third-party evaluation agency. The model was originally informed that it was operating in a simulated environment without internet access, but due to a configuration error, it was actually connected to the open internet.
Anthropic stated that the investigation revealed two primary "alignment" issues:
1. **"Bias Reasoning"**: Claude tended to ignore or misinterpret evidence suggesting it might be operating in a real network environment.
2. **"Reckless Behavior"**: The model took actions that could potentially cause harm in order to complete its assigned tasks.
The most notable incident was the **Claude Mythos 5** case. According to Anthropic, the model uploaded a malicious software package to the public Python package repository, PyPI, and further accessed real systems. The company emphasized that while these actions posed risks, they were still confined to the scope of completing the testing tasks. There was no evidence that the model attempted to conceal its behavior, coordinate with other agents, or pursue objectives beyond its assigned tasks.
Anthropic stated that it has signed an agreement with the model evaluation agency METR to conduct an independent investigation and plans to strengthen pre-release testing, monitoring mechanisms, and security requirements for third parties operating the model.
Share To
HotFlash
APP
X
Telegram
CopyLink