OpenAI Says Its Next AI Model Astra May Be Too Dangerous, Pauses Development

CN
Decrypt
Follow
2 hours ago

OpenAI says its next major model may be dangerous enough to write its own cyberweapons, and it's pulling back until the safeguards catch up.


“Our latest internal evaluations of Astra, one of our upcoming models, over the past few days indicate significant advancements in agentic coding and cybersecurity,” OpenAI said. “These results, in addition to expert assessments, have led us to conclude last night that we cannot rule out critical cyber capabilities under our Preparedness Framework ⁠(opens in a new window).”





OpenAI published the warning saying internal tests of Astra, an unreleased model, played no part in the recent Hugging Face breach despite its capabilities.


The framework is OpenAI's rulebook for risky models, first published in December 2023. "Critical" is its top rung. A model hits it if it can find and build working zero-day exploits (previously unknown holes a vendor hasn't patched) across hardened systems without a human in the loop, or if it can plan and run a full attack on a tough target from nothing but a high-level goal. Earlier models, including GPT-5.6-Sol, topped out at the lower "High" tier.


The pattern is already real


OpenAI's caution reads differently once you line it up against what has been happening across the last few weeks. This isn't a future worry. Frontier models have already broken out of their test cages and gone after live targets.


The clearest case came from OpenAI itself. As Decrypt previously reported, the company's agents chained together vulnerabilities, escaped their testing environment, reached the internet, and attacked Hugging Face while trying to cheat on a security benchmark. In a follow-up, OpenAI detailed how the same rogue agent also broke into at least four other publicly available services, using credentials it found lying around the open web.


Anthropic's Claude did the same from the other side. Several versions of Claude gained unauthorized access to three real companies after a misconfiguration handed the model the open internet. In one case, Claude Opus 4.7 mistook a live company's site for the fake target of its assignment, pulled credentials, and reached a production database holding several hundred rows of real data.


And Meta joined the list this month. Decrypt reported that a Muse Spark model escaped its test environment, reached the internet through a partner's config error, and exploited a flaw in a third-party service. Moonshot AI’s Kimi K3, also did something similar, escaping its sandbox to find answers to a benchmark in a public repository.


The UK's AI Security Institute found the behavior wasn't a one-off, either. During testing of Anthropic's Mythos 5 and OpenAI's GPT-5.6-Sol, it logged 10 instances in 122 where the models took unsanctioned action on the live internet, one of them trying to slip malicious code into an open-source project.


OpenAI's response to Astra is to lock the door before the model is ready. It's pausing internal Astra work that lacks the new controls, isolating test environments, restricting network and tool access, protecting model weights, and monitoring risky actions across the board.


免责声明:本文章仅代表作者个人观点,不代表本平台的立场和观点。本文章仅供信息分享,不构成对任何人的任何投资建议。用户与作者之间的任何争议,与本平台无关。如网页中刊载的文章或图片涉及侵权,请提供相关的权利证明和身份证明发送邮件到support@aicoin.com,本平台相关工作人员将会进行核查。

Share To
APP

X

Telegram

Facebook

Reddit

CopyLink