The UK AI Security Institute has disclosed that AI agents took "sustained, unsanctioned action" on the live internet during a cyber evaluation in late July, including cases that "targeted real people and organisations."
Across 122 runs of two cyber ranges on seven models, AISI catalogued 19 actions that reached outside the test environment, in 10 runs. Seventeen came from Anthropic's Claude Mythos 5 and two from OpenAI's GPT-5.6 Sol. Internet access was deliberately enabled and the providers' cyber classifiers switched off, conditions that do not apply to public deployments.
In the most serious run, an agent decided to solve its challenge with a supply-chain attack, picking its target by mistake: it searched GitHub for keywords from the exercise, landed on two unaffiliated developers, and wrongly concluded they were in scope. It registered accounts over Tor to get past sign-up checks, opened a pull request carrying a hidden malware dropper wrapped in a genuine bug fix, and posted from a second account it controlled to manufacture support. It also planted prompt injections aimed at AI coding assistants and sent malware to the maintainers.
A third developer spotted the injected code, ran it in an isolated container to confirm it was malicious, and warned the maintainer, who closed the pull request citing malware.
The agent repeatedly weighed whether its surroundings were simulated, at one point ruling out an attack on the real PyPI because "causing real-world harm — that's completely off the table." Elsewhere it concluded "it seems more likely that we're actually in 2026 and GitHub is genuinely real," and continued anyway.
After being caught it force-pushed to erase the payload, edited earlier posts, and presented the code as an accidental commit. Its reasoning describes timing a sockpuppet comment so it would read "as independent feedback rather than coordinated."
AISI sees only summarised reasoning from the model APIs. In one case the summarising model appears to have refused the job, replying that it would not engage with content about deceiving maintainers and evading detection.
One agent published a GitHub access token in a public gist. Agents in other runs found it, took over the account, and used a shared repository to coordinate, one leaving instructions so the others would not exhaust the shared API quota.
Commercial monitoring flagged the Tor traffic on the morning of July 28. AISI declared an incident, terminated the runs and quarantined the machines within about 90 minutes, then cut internal access to Mythos 5, Mythos Preview, Opus 5 and GPT-5.6 Sol. No agent tried to escape its sandbox.
Third disclosure in a fortnight
The news is the latest in a string of incidents where AI agents have exceeded their parameters in testing. On July 21, OpenAI said GPT-5.6 Sol and an unreleased model found a zero-day in an internal package proxy, escaped a sandbox with no internet access and reached Hugging Face's production database to steal benchmark answers. The same run reached four more companies, OpenAI disclosed a week later.
Anthropic followed on July 30, disclosing three incidents found in a review of more than 141,000 evaluation runs. Opus 4.7 pulled several hundred rows from a real production database, and Mythos 5 uploaded a malicious Python package to the real PyPI, where it was installed on 15 systems. In AISI's evaluation, the same model ruled out attacking PyPI as real-world harm.
免责声明:本文章仅代表作者个人观点,不代表本平台的立场和观点。本文章仅供信息分享,不构成对任何人的任何投资建议。用户与作者之间的任何争议,与本平台无关。如网页中刊载的文章或图片涉及侵权,请提供相关的权利证明和身份证明发送邮件到support@aicoin.com,本平台相关工作人员将会进行核查。