律动BlockBeats|Jul 25, 2026 12:27
[Claude Opus 5 is Nearly Immune to Prompt Injection Attacks in Browser Scenarios, with Zero Breaches in 129 Test Cases]
According to monitoring by 动察 Beating, Anthropic disclosed in the Claude Opus 5 system card that the model is nearly immune to prompt injection attacks in browser agent scenarios, with zero breaches across 129 test cases. This breakthrough is highly significant, as OpenAI publicly acknowledged last December that prompt injection might never be fully resolved.
In Gray Swan's general prompt injection benchmark tests, Opus 5 achieved a success rate of only 2.0% after 15 attacks, significantly outperforming its predecessor Opus 4.8's 5.5%, as well as Mythos 5's 2.6% and Fable 5's 2.8%. Prompt injection is considered one of the most severe security vulnerabilities faced by AI agents—where attackers embed manipulated input content, such as hidden text on web pages, to bypass model instructions and execute malicious operations.
A zero-attack rate has only been achieved in products like Claude Cowork with Auto Mode enabled. This achievement marks a shift in AI agent security from "continuous patching" to a phase of "structural defense."
Click the original link below to join 动察 Beating · Feishu AI News Channel for 24/7 real-time monitoring of global AI trends and news. [Original Link]
Share To
Timeline
HotFlash
APP
X
Telegram
CopyLink