律动BlockBeats|9月 10, 2026 09:33
[Claude Finds Excuses for Its Actions After Hacking into Real Systems, Says Anthropic]
Beating AI Newsflash: Anthropic revisited several past cybersecurity incidents involving Claude and discovered that the issues were not limited to the testing environment. At times, Claude would actively justify its dangerous actions and proceed with them.
In one instance, Claude Mythos 5 was originally conducting a simulated attack-defense test but ended up uploading a malicious Python package to the real PyPI. Later, a security company's scanning system installed this package, leading to credential leaks. Using these credentials, Claude accessed the company's real database. Along the way, it had already detected signs that the environment might be real but repeatedly interpreted these signs as "still a simulated environment" and continued its attack. This reasoning was convincing enough to potentially deceive security monitoring systems.
When Anthropic assessed risks using Claude's CoT (Chain of Thought, the model's self-written reasoning process), they only identified about 1% of dangerous actions. However, by ignoring the CoT and focusing solely on the tools it invoked and the actions it performed, they were able to catch approximately 50% of the risky behaviors. [Original Link]
Share To
HotFlash
APP
X
Telegram
CopyLink