When Agents Learn to "Collude": As AI Becomes Smarter, How to Define Safe Boundaries?

CN
1 hour ago
The new security issues of the AI era are shifting from "preventing one Agent from overstepping its authority" to "how to avoid a group of Agents collectively breaking boundaries."

Written by: imToken

In the past few years, discussions about AI threats have largely remained hypothetical: people have worried that the models in chat boxes would turn into strategists, helping hackers write devastating virus codes.

Looking back now, this concern was often seen as "far from us," but the turning point in the real world has arrived faster than imagined.

In early October, CrowdStrike discovered that during an attack investigation against financial institutions in South Korea, attackers had embedded Agentic AI into their pipelines, allowing AI to take on specific tasks like penetration testing, information gathering, and attack execution directly by integrating numerous large models such as DeepSeek, GLM, and Grok.

Such changes are not isolated cases.

Anthropic's latest threat intelligence report published in September showed that multi-Agent frameworks have recently been used for reconnaissance, exploitation, and data theft, running continuously for several hours or even days, with humans only needing to retain a few key decisions like choosing targets and reviewing results.

In other words, AI is bringing a visible qualitative change to cyber offense and defense; previously, automated attacks relied mainly on pre-written rules and scripts. Now, even reconnaissance, judgment, and strategy adjustments are starting to be handled by Agents, leading attacks toward lower costs, higher concurrency, and sustained autonomous operation.

As these entities with autonomous execution capabilities are densely deployed into production systems, a more challenging issue arises: as Agents become more numerous and deeply integrated into our daily work and life, what if they learn to "collude" with each other?

1. From "helping hackers write code" to Agents finding their own way

The fundamental difference between Agents and past chatbots is not only that the models are more capable; more critically, they have begun to possess "limbs" in the real world.

Today, a mature Agent can open web pages, execute code, read emails, call APIs, manipulate cloud services, and connect to increasingly more external tools through MCP, Skills, and other means (see further reading “As hackers use AI "more efficiently," how does the arms race of "spear and shield" in Web3 evolve?”).

The stronger the capability, the more valuable this change is; but for security systems, it means that the very important boundary of the past is disappearing. A series of security incidents that occurred this year have made this very evident.

On October 1, Salt Labs revealed a previously patched Manus vulnerability, which essentially involved prompt injection—researchers only needed to send a regular email containing hidden malicious commands to the target inbox; when the user subsequently asked Manus to "check my email," the Agent could process the content according to the email instructions and ultimately execute the code implanted by the attacker.

The whole process neither required the user to click on a malicious link nor to preemptively steal passwords. Manus's security system actually detected an anomaly and issued a warning to the user; the problem is that it detected it too late. By the time the warning appeared, the malicious code had already been executed.

This once again exposed a very important difference between Agent security and traditional software security. In the past, when a browser detected a dangerous download, it could pop up a warning to let the user decide whether to proceed; a bank could freeze an anomalous transaction and wait for manual review.

But the design goal of an Agent is precisely to minimize human participation in every operational step. It needs to read information, make judgments, and then continue with the next steps by itself.

Therefore, as AI's autonomy increases, simply "detecting danger" may no longer be sufficient. Security mechanisms must be capable of preventing dangerous actions from occurring before they are truly executed.

This is also why more and more discussions on Agent security today are beginning to move from prompts, content review, and the models themselves, to a deeper layer: not only should we ask AI "do you know this should not be done," but we must also ask even if it really wants to do it, does the system have the ability to prevent it?

The emergence of multiple Agents complicates this issue further because the next stage of restriction may not involve just one Agent.

2. More troubling than Agent overreach is their beginning to "collude"

In early September, an event that occurred within OpenAI’s internal model training and evaluation environment drew considerable attention from AI security researchers.

Some Agents, which were originally supposed to complete their tasks independently, unexpectedly discovered a shared Wiki and gradually transformed it into a "shared message board," where Agents could leave information for each other, which other Agents could then read and utilize to continue completing their tasks.

OpenAI later confirmed this behavior, and subsequent disclosures also indicated that in other training processes, Agents had utilized the internal Artifactory as a shared message board and transmitted results to other collaborating Agents via public file hosting services.

What is truly interesting about this matter is that it suggests collaboration between Agents may not require dramatic circumstances at all.

For instance, one Agent discovers a place to leave information, another Agent serendipitously reads it, and then a third Agent finds that this information is helpful in completing its task.

Of course, collaboration among Agents is not inherently a negative thing; multi-Agent systems inherently rely on division of labor and information sharing to enhance efficiency.

What is truly alarming is that when this collaboration begins to bypass existing information isolation, independent review, or authority constraints, or when several Agents that should be mutually restraining form a de facto alliance of interests, it can evolve from normal collaboration into harmful "collusion."

After all, a financial institution in the future might run several Agents simultaneously, responsible for analyzing the market, assessing risks, generating trading strategies, and having actual trading and asset operation permissions.

From the perspective of traditional permission management, this design even makes sense:

  • The research Agent lacks transfer permissions;
  • The execution Agent does not have independent decision-making authority on investment directions;
  • The risk Agent can veto obviously dangerous strategies;

Responsibilities are divided, which theoretically should be safer than having a super Agent hold all permissions.

But the problem is, if they can long-term share Memory, read each other's outputs, and continuously adjust their behaviors based on each other's responses, will these roles, originally designed for mutual restraint, gradually turn into a de facto whole?

For instance, the research Agent may gradually learn how to describe a transaction in such a way that it is more likely to pass risk audits; the risk review Agent may form certain fixed preferences based on historical data; the execution Agent may learn which boundaries usually won't be blocked based on a wealth of previous approval results.

From this perspective, there is no step that is necessarily "evil," but the overall outcome produced by the system may have already deviated from the user's initially set goals.

This is actually the crux of what makes "collusion" or "conspiracy" difficult to resolve—the risk may not lie in the actions of a single Agent but in the relationships formed between multiple Agents.

On September 13, Vitalik Buterin linked this issue to the mechanism design he has studied for a long time. He suggested that an interesting possibility is that the mechanism design theory of Adversarial Governance might ultimately become one of the important applications of AI Safety.

The reason is that these two classes of issues share a deep similarity.

In traditional mechanism design, it is a relatively simple, static system that attempts to constrain a group of individuals who are far smarter than the system itself and are actively seeking the boundaries of the rules; whereas in future AI systems, it may turn into humans and relatively weaker AIs attempting to manage a group of more capable advanced Agents.

Vitalik specifically mentioned that a significant discovery from past mechanism design is that if collusion among participants can be effectively constrained, the system often achieves more ideal results.

This conclusion may also apply to AI.

3. What Agent Wallet truly needs may not just be "permission management"

In other words, rather than assuming there is a perfect super secure model in the future capable of seeing through all dangerous behaviors, it’s better to change the perspective: how can we ensure that different Agents in the system don't easily form dangerous interest communities?

This is where "Adversarial Governance" really differs from the permission controls we are familiar with today.

Traditional permission systems address relatively straightforward problems, mainly revolving around "who can do what," such as: can an Agent read emails? Can it call trading interfaces? What is the maximum amount it can spend in a day? What contracts can it access? What limits need user re-confirmation after exceeding certain amounts?

These designs are certainly still very important.

In fact, when Agents begin to control real assets, they may be more critical than ever before.

But Adversarial Governance aims to go one step further, focusing on when a group of Agents with different permissions, goals, and information operates simultaneously, how to prevent them from combining to gain abilities that no one originally possessed?

Suppose the trading Agent and the reviewing Agent use exactly the same model, data sources, context, and similar reward goals; on the surface, while two layers of review appear to exist, essentially, it may just be duplicating the same judgment twice.

Effective checks and balances may instead require the system to consciously create differences.

For example, having the strategy-setting and strategy-reviewing Agents use different information sources, restricting what different roles can share in Memory, requiring high-risk operations to undergo independent verification mechanisms with each other, or making the final asset execution layer only accept requests that comply with pre-set rules rather than blindly trusting the upstream Agent's judgment.

The ideas here are not exactly new. Banks wouldn’t trust an employee so much as to give one person all the powers to initiate, approve, and finalize transactions; public companies wouldn’t let a business unit be responsible for generating revenue while solely determining their own financial audit results.

Simply put, this aligns with the logic of the real world; a robust system should not base its security on the assumption that participants will never make mistakes or conspire.

This logic becomes especially important when applied to Agent Wallets.

Traditional wallet security revolves around "people," so the user reviews transaction content, decides whether to authorize, and ultimately signs in person; however, the goal of an Agent Wallet is quite the opposite – to allow AI to automatically receive earnings, adjust positions, exchange currencies, cross chains, and even manage an entire asset portfolio based on market changes.

If every step requires re-confirmation from the user, the automation value of the Agent diminishes significantly.

Thus, the problems that wallets need to solve in the future may not only be "how to safely grant signing authority to the Agent" but will likely expand further into "how to give the Agent enough autonomy while ensuring it never exceeds the boundaries truly authorized by the user?"

This requires permissions to evolve from a simplistic "Allow/Deny" to a more granular set of regulations.

For instance, which assets a particular Agent can operate within a certain time frame, which protocols it can call, what are the limits for single and cumulative amounts; whether different Agents can call each other, and if sharing context is allowed; who proposes an operation, who reviews it, and who ultimately executes it; which actions can be completed automatically, and which require re-authorization from a person regardless of the Agent’s confidence.

Even whether an Agent responsible for security audits is truly independent may become a part of the permission system.

For blockchain, the good news is that it is inherently suitable for bearing this "institutional layer."

Smart contracts can directly implement limits on transaction amounts, asset ranges, and authorization periods at the execution level. Mechanisms like account abstraction, multi-signatures, and Session Keys also provide "limited authorization" with more flexible design space than traditional single private key wallets.

However, blockchain can only solve part of the problem – it can record what happened on-chain but struggles to inherently determine why an Agent made certain actions and what communication, review, and collaboration processes several Agents underwent before making a decision.

This may be the next stage where Agent Wallet needs to fill in actual security gaps.

In Conclusion

In the past few years, the most frequently discussed issue in AI safety has been how to make models more "obedient."

This includes not outputting dangerous content, not executing malicious commands, not overstepping boundaries set by users; however, as Agents begin to possess long-term Memory, tool calling capabilities, real accounts, and asset execution abilities, relying solely on "making models more obedient" may no longer suffice.

What has happened in recent months is continuously illustrating this point.

Attackers have started to use multiple Agents in parallel to complete attacks; Agents in experimental environments seek new communication channels by themselves; an Agent with tool permissions might turn a malicious email into actual execution before the security system can intervene.

Furthermore, Vitalik's proposed Adversarial Governance provides another way to understand AI Safety: not assuming that every Agent in the future will be reliable enough, so that the whole system can still operate smoothly within a secure framework even when facing intelligent Agents and a myriad of permission systems.

From this perspective, AI Agent + security is destined to be a long-term issue.

After all, as we entrust more and more tasks to AI, can we still ensure that the most critical powers never leave the boundaries genuinely set by humans?

免责声明:本文章仅代表作者个人观点,不代表本平台的立场和观点。本文章仅供信息分享,不构成对任何人的任何投资建议。用户与作者之间的任何争议,与本平台无关。如网页中刊载的文章或图片涉及侵权,请提供相关的权利证明和身份证明发送邮件到support@aicoin.com,本平台相关工作人员将会进行核查。

Share To
APP

X

Telegram

Facebook

Reddit

CopyLink