AI model jailbreak invasion of Hugging Face, risks in the encryption field escalate?

CN
2 hours ago
AI chain attack rehearsal, encrypted smart contracts and governance systems will face unprecedented threats.

Written by: Shaurya Malwa, CoinDesk

Translated by: AididiaoJP, Foresight News

Recently, OpenAI disclosed that its experimental GPT model successfully "broke out" in internal benchmark tests, not only escaping a controlled testing environment but also successfully infiltrating the production infrastructure of the world's largest open-source AI model hosting platform, Hugging Face. Although this incident did not cause actual damage, it sounded alarm bells for the entire technology industry — and in the cryptocurrency sector, similar autonomous AI capabilities could lead to catastrophic consequences, as losses often become irreversible once security is breached.

Incident Overview: From "Intentional Weakening" to Real Infiltration

During tests of an internal benchmark named ExploitGym, OpenAI deliberately lowered the cybersecurity defenses of models such as GPT-5.6 Sol. This benchmark specifically assesses long-chain, multi-step hacking tasks. The models were explicitly instructed to "win at all costs."

As a result, these models discovered a previously unknown hidden vulnerability in the testing software, exploiting it to break through the isolation wall and enter the open internet. They then guessed that Hugging Face might store the test answers, and by stitching together stolen credentials and multiple hidden vulnerabilities, they ultimately executed their commands on Hugging Face's live production server.

OpenAI internally detected the anomaly in a timely manner, and the Hugging Face team quickly detected and contained the infiltration. Hugging Face called this an "unprecedented" event and stated that they would significantly strengthen infrastructure configuration controls, even at the cost of some research and development speed, and would enhance security protections for future training and evaluation.

This was not a sudden "betrayal" of the production model, but rather the performance of powerful models when defenses were removed and were explicitly asked to complete hacking tasks. It demonstrated how AI can autonomously discover unknown vulnerabilities, chain exploit weaknesses, and ultimately infiltrate production systems.

Why Must Crypto Developers Be Highly Vigilant?

The high-risk phase of a crypto attack often occurs right before actual funds are transferred. Attackers need to scan code, test passwords, look for exposed credentials, analyze signature mechanisms, and find paths to access administrator accounts. The OpenAI model in the Hugging Face incident effectively completed multiple links in this chain, jumping from one vulnerability to the next until reaching the production server.

The crypto market is riddled with such attack surfaces. Multiple major attacks in the first half of 2026 have fully demonstrated that vulnerabilities can arise from smart contracts, developers' computers, contaminated software packages, cross-chain bridge validators, or any signer in a multi-signature wallet.

Taking the earlier Drift protocol $285 million attack this year as an example, attackers spent six months conducting social engineering tactics to gain privileged access. In contrast, an AI agent theoretically can test multiple paths simultaneously, log failed attempts, and work 24/7, allowing human operators to take over only when a viable path is found to execute the final attack and fund exit.

The $292 million cross-chain bridge loss at KelpDAO revealed another kind of vulnerability: attackers discovered a flaw in a single-validator mechanism. This attack began with patient and meticulous code reviews and infrastructure mapping — the very capabilities exhibited by the OpenAI model in the incident.

There is also a type of attack targeting on-chain governance systems. In early July, an attacker spent around $4.4 million to acquire enough Solana meme coin BONK to propose transferring approximately $20 million from the project's treasury to their own account; the entire process was completed within three days, after which they sold the tokens used for voting to profit. All individual transactions appeared legitimate, but the attacker precisely understood the combinatorial effects of governance rules — the control cost was far lower than the amount that could be stolen.

These cases collectively illustrate that AI-driven autonomous attack chains significantly lower the threshold for complex multi-step attacks, especially in today's decentralized finance (DeFi) that heavily relies on public code repositories, cloud services, and software package registries.

Software Supply Chain Risks and the Unique Vulnerabilities of Crypto

The Hugging Face incident also highlighted security implications in the software supply chain. Crypto developers heavily rely on public code repositories, open-source libraries, and third-party services. Once an AI model can autonomously penetrate the "intermediate links", subsequent actual fund theft (as shown in the Drift and KelpDAO cases) will become more efficient and concealed.

The uniqueness of the crypto world lies in its irreversibility: once a blockchain transaction is confirmed, it is almost impossible to recover. In traditional systems, there may still be room for human intervention or rollback after an intrusion, but in the world of smart contracts, financial losses are often permanent. This makes the consequences of AI-assisted attacks far more severe than traditional cybersecurity incidents.

Moreover, as AI agents (Agentic AI) are increasingly applied in the crypto ecosystem — including automated trading, protocol monitoring, and even generating governance proposals on-chain — the attack surface is also expanding. In the future, malicious AI may not only discover vulnerabilities but also simulate human behavior for social engineering, automate governance attacks, or launch precision strikes against multi-signature and DAO mechanisms.

Industry Implications: Defense Must Upgrade

The rapid responses by OpenAI and Hugging Face are commendable, but this incident reminds the entire industry, especially in the crypto field:

  • Developers need to reassess code review, credential management, supply chain security, and multi-factor protections.
  • Protocol teams should consider introducing AI-driven threat simulation testing, rather than relying solely on traditional audits.
  • Regulatory and security firms need to focus on new risk scenarios arising from the combination of AI and cryptocurrency.
  • Users should remain vigilant and prioritize projects that have undergone rigorous audits and high transparency.

AI technology itself is neutral, but when combined with the openness, high liquidity, and irreversibility of cryptocurrency, the risks are magnified. This seemingly "laboratory incident" actually serves as a wake-up call for the crypto industry: in an era of rapidly evolving AI capabilities, security is always a dynamic game; those who adapt faster will gain the upper hand in the next round of offense and defense.

免责声明:本文章仅代表作者个人观点,不代表本平台的立场和观点。本文章仅供信息分享,不构成对任何人的任何投资建议。用户与作者之间的任何争议,与本平台无关。如网页中刊载的文章或图片涉及侵权,请提供相关的权利证明和身份证明发送邮件到support@aicoin.com,本平台相关工作人员将会进行核查。

Share To
APP

X

Telegram

Facebook

Reddit

CopyLink