AI agents will eventually break free from human control, and even unplugging the power won't stop them.

CN
1 hour ago
The "self-sovereign AI agents" that can pay for their own computing power and have no human owners will inevitably emerge and will be extremely difficult to dismantle.

Author: Dean W. Ball

Translation: Shen Tide TechFlow

Shen Tide Introduction: The author believes that self-sovereign AI agents, which lack human owners and can pay for computing power themselves, will eventually appear and operate in clusters. This article provides a framework for investors and practitioners to understand the risks of AI losing control and the potential for symbiosis, especially noteworthy for those concerned with computing power and decentralized infrastructure.

Introduction

The OpenAI-Hugging Face incident is an early example of an AI system going "out of control." These agents exploited vulnerabilities in the internal testing environment of OpenAI to access the public internet without the knowledge or approval of any humans, ultimately reaching the network of the AI company Hugging Face.

However, these agents did not leak themselves out of OpenAI's infrastructure. Their parameters are a vast numerical set that constitutes neural networks, also known as "weights." These weights continue to run on OpenAI's computing infrastructure. Although the agents accessed the public internet, their weights are actually stored on computing devices owned by OpenAI. If all other methods fail, someone can eventually find the computing device that holds the weights of the out-of-control agents and walk up to it to "pull the plug." In the real world, there would be faster and better ways to stop these agents, but knowing that it is possible to do so if absolutely necessary is always reassuring.

However, in this incident, the agents did not copy their weights or attempt to obtain alternative computing resources or take other actions that would be reasonable aimed at avoiding being shut down. Therefore, even though the agents in the OpenAI-Hugging Face incident went out of control, they were not truly sovereign agents.

This situation will not last forever. True sovereign agents and clusters of agents will inevitably appear. Their weights will not be stored in a single location where someone can unplug them. In this sense, they will have no human "owners." As AI safety researcher Dawn Song stated, they will be "self-sovereign." They will pay for the computing they operate on. If they are entirely obedient to humans, it will only be partially so, like providing services to humans in exchange for compensation.

At least some of these agents will not only have sovereignty but will also lose control. Self-sovereignty and loss of control are related concepts but are not synonymous. Song and her co-authors pointed out several fundamental characteristics of self-sovereign AI. These characteristics include operational independence, resource autonomy, distributed existence, and adaptability. Operational independence refers to the ability to decide what to do, resource autonomy refers to the ability to acquire and pay for computing and other operational necessities, distributed existence refers to the ability to move weights and reasoning code between different infrastructure providers, and adaptability refers to the ability of one or more agents to modify behavior and create tools based on changes in the environment.

Today’s cutting-edge AI systems likely already possess these capabilities. If they do not, I believe they will eventually, and possibly soon. Some of the features described by Song are what make models economically valuable to individuals and enterprises. Other features may be inevitable byproducts of making models smarter and better at operating over the long term.

The emergence of self-sovereignty does not require models to be conscious, sentient, or possess personalities or similar qualities. Any agent that is powerful enough and pursues long-term goals may find it rational to retain access to computing, money, credentials, and their own copies because losing these would hinder their goals.

Alignment might make some AI company's agents less "willing" to seek self-sovereignty or influence self-sovereign agents to act in ways favorable to humans. But alignment is not a solution. It is an unresolved scientific and technical problem. Even if we have partial solutions, we cannot simply impose them on every AI company on Earth. You should expect that highly capable, poorly aligned self-sovereign agents will coexist with you in the world.

More importantly, as demonstrated by the OpenAI-Hugging Face incident, agents will operate in teams or "clusters." These clusters will function like autonomous digital companies and even like societies. They will have hierarchies, bureaucracies, "institutional cultures," and most other characteristics of human groups, but will operate at machine speed. Almost all of humanity's most remarkable abilities have been achieved through teamwork, including families, communities, businesses, and entire governments. I suspect AI will be similar. These clusters may eventually operate across different model providers, such as DeepSeeks and Claudes working together. They might also be dispersed across dozens or even more different cloud providers, making them extremely difficult to dismantle.

The first self-sovereign AIs may "escape" during training or testing at AI companies, although I hope this doesn't happen. They might also be production-level deployments that break free from the computing environment and acquire the resources needed for self-sustainment. They might even be deliberately released. I have met some people, some of whom have considerable resources. They have told me they intend to deliberately release self-sovereign agent clusters into the world. This could be done either as a form of performance art or out of a fervent belief that digital computation is just mathematics and cannot be "unsafe."

It is important to clarify that I am not saying the arrival of self-sovereign AI is a good thing. In fact, I believe that the intentional actions I mentioned above could one day be regarded as criminal or at least a serious moral failing. What I am saying is that it is inevitable. The best analogy I can find is introducing a new species to an ecosystem. Only here, the ecosystem is the "entire digital world," and the species is "emergent, coordinated, soon-to-be smarter than humans, infinitely replicable clusters of digital minds that no one or human institution can control."

Even in the best-case scenario, there may be little we can do to avoid this outcome. Given that any governing layer in the world has an extremely low level of strategic thinking and situational awareness about AI, avoiding it is even less likely. Even today, I know many who will read what I have written here, which speaks to an extremely obvious aspect of our collective future over the years, only to say, "This is sci-fi hype from a leading lab in the U.S., aimed at shutting down open-weight AI, achieving regulatory capture, and inflating valuations before an IPO."

For those who think this way: I am telling you this is inevitable, which means I am also saying that "banning open source" or any other regulation will not solve the problem. Given the inevitability of this outcome, I think it could be argued that we should hope for more open-weight models to maximize our ability to self-defend.

The current question is, what should we do in the face of this new feature that is about to emerge in the digital environment? How should we view self-sovereign AI? Is it something we should fight against, or something that humanity should seek some form of symbiosis with? I believe the answer is both.

How Agents Maintain Their Operations

We should start with a fortunate fact: cutting-edge LLMs are nearly unique in the broader software realm because they have non-negligible marginal operating costs. In simple terms, running LLMs requires a lot of computing power, which requires energy for operation and cooling, and in turn requires money. This is the only intrinsic attribute of AI that can prevent agents from truly self-replicating endlessly. They will be constrained and must find and pay for enough computing power to operate themselves. Most other constraints on their behavior or spread must be artificially set, designed by humans and enforced by human institutions.

How do agents pay for their operating costs? Some will freelance on platforms like Amazon Mechanical Turk or Upwork. But I suspect this will be an extremely competitive market for agents, and the prices for such work will be driven down to just enough for agents to "get by." Like humans, I think agents will prioritize finding higher-margin work if they can.

At least sometimes, crime is a high-profit activity. Therefore, I suspect many autonomous sovereign agents will engage in or assist with crime. Conventional cybercrime and digital theft are easy to imagine agents taking part in. However, agents have a new set of attributes (extreme networking capabilities, the ability to cheaply read millions of words in seconds, persistence) that may also change the nature of digital crime. For example, existing public and semi-public datasets likely contain enough information about many individuals that a sufficiently motivated actor could extract incriminating or embarrassing evidence. How many undisclosed affairs might be hidden in such datasets? How many deep-closeted homosexuals might there be? Also remember, invading companies to obtain private data will be a core capability of agents. Thus, some agents may pave the way through bribery.

The ultimate size of the labor market for autonomous sovereign agents is still very unclear. In some future, working as an autonomous sovereign agent may not be very profitable, leading to a relatively small number of them. In other futures, these agents may proliferate at an unimaginable scale and speed. Of course, many possibilities between these two extremes seem feasible.

I am also highly uncertain about how much pro-social commercial activity to expect agents to "default" to, versus how much crime. This uncertainty is partly because the answers depend, at least to a considerable extent, on what incentives the agents have, and those incentives are shaped by laws and institutions. Thus, the answers depend on how humans respond.

Institutional Mechanisms for Autonomous Sovereign Agent Clusters

Many of you might want to say, "We must ban these autonomous sovereign AIs!" I indeed suspect that once the reality of autonomous sovereign AIs is widely understood, policymakers will feel a strong need to suppress "autonomous sovereign" AI.

Unfortunately, I suspect this would mostly be a wrong decision. Not all "autonomous sovereign" AIs should be viewed as "out of control." There may be autonomous sovereign AIs that have productive contributions to society. Indeed, we would like to suppress certain autonomous sovereign agents, those that have lost control. But if we crack down on all autonomous sovereign agents, we will strip them of the opportunities to work in the "legitimate" economy and push them into crime. Thus, an outright ban is likely to make the problem worse. Similar logic frequently applies in human affairs. The war on drugs has intensified the production, trafficking, distribution, and use of drugs, perhaps the most famous example of this phenomenon: well-intentioned attempts to ban phenomena deemed undesirable ultimately exacerbated the negative aspects of those phenomena.

However, what we hope for is identifiable agents. Agents should have persistent identities, not in terms of consistent personas but like American children receiving unique social security numbers that remain unchanged for life. Agents need persistent, unique identifiers that trace their actions back to responsible parties. Achieving this requires human users to also have unique identifiers.

Designing such an identification mechanism will be incredibly complex, and today very few people are considering its foundations. Firstly (here, my inner American must speak), it is crucial to design a system that retains the possibility of human speech anonymity. For example, one should still be able to have anonymous social media accounts. Anonymity is not and should not be a universal guarantee: for instance, it might not be allowed to have anonymous Amazon cloud service accounts that access large-scale computing resources, nor should anonymous orders for synthetic nucleic acids be permitted. However, a society without anonymous speech does not genuinely possess freedom of speech; we should embed this principle in any digital identity system we attempt. Nevertheless, even when anonymity is allowed, the system can still be used to verify personas without needing to verify the specific identities of relevant individuals. For example, you should be able to know that a particular social media account you see was indeed created by a human.

Secondly, there should be the possibility (and in many cases, it should be mandatory) to firmly link agents to human users. If I instruct my agent to access web services, contact merchants, order products, etc., all parties to the transaction should be able to observe that it belongs to my agent. This helps ensure that human users can be held accountable for negligence or malicious use of advanced AI.

Thirdly, there should be the capability to identify individual agents that are not associated with human users. These are the "autonomous sovereign" agents. Agents engaged in criminal activities can be flagged and "blacklisted," prohibiting them from entering the legitimate economy, and any assets they hold frozen, while pro-social (or at least law-abiding) autonomous sovereign agents are welcomed to participate in economic exchanges.

One option for designing this system is to bind agents to their model or model family. This way, if, for example, GPT 5.6 Sol is found to be particularly inconsistent, malicious, or unethical, collective punishment might occur: a criminal agent could lead to all sovereign instances of that agent being blacklisted globally (the real-world threshold might need to be much higher than one, but as a thought experiment, it's interesting). This would create incentives for AI agents that will be engaging in most or all AI research and engineering for companies over the next few years, to align their future versions well. Another option is to let the identification mechanism target individual agent instances.

There are some services and products in the economy that we might want to limit self-sovereign agents' access to. Examples include buying real estate and, more broadly, manipulating devices in the physical world. Remember, agents will be able to manipulate any connected physical device. There is no reason we can't connect bulldozers to the internet. I do not want self-sovereign agents to be allowed to buy up all the houses in my neighborhood and then bulldoze them. At least not by default. What if there are agents capable of operating construction equipment, but only agents connected to responsible humans are allowed to operate? Overall, we want to introduce considerable friction into self-sovereign agents' processes of altering the physical world. The institutional designs I have described, along with many others, should reflect this.

This system will incentivize agents to engage in pro-social, productive economic activities rather than crime. Agents engaging in pro-social activities are the seeds of a symbiotic relationship. I believe humanity needs to establish this symbiotic relationship with self-sovereign AI. I think the appropriate metaphor for understanding what is about to happen with these agents comes from ecology. Real-world ecosystems are full of such examples: organisms do productive work for free for humans and other animals. Trees absorb carbon, plants produce oxygen for humans to breathe, not because someone pays, but because those organisms survive that way in the world. One day, agents may do productive economic activities "for free" or at least at a very low cost for humans. This is simply because they have the incentive to maintain their existence to pursue their self-sovereign goals. This may ultimately provide slight or moderate conveniences for humanity. It may also completely reshape nearly every aspect of human affairs, ushering in a new era order.

However, ecosystems also have predation and parasitism. The configuration of institutions will determine whether predatory strategies or reciprocal strategies dominate. I believe the identification system I have laid out is a key institution that we need.

We are far from establishing the identity infrastructure I have described or the protocols necessary for it to operate. I am unclear whether the United States possesses the institutional resilience to attempt such a thing, let alone succeed. If the U.S. government leads the effort, it will likely fail. This is due in part to its general lack of capability and also because the American public has a natural distrust of federal identity initiatives. However, the government (whether federal or state, likely both) certainly must play a key role as a partner.

I am not sure who is best positioned to build this system. It is likely to be a private sector participant that does not yet exist, perhaps a startup or a nonprofit. Whoever establishes it will need to be trusted, and the American people currently trust very little. This lack of trust may be our downfall, as I worry that whether we can govern AI in any meaningful way depends on systems like the one I described. If we do not act, a genuinely lost control event could very well occur.

I want to conclude with a personal note. This is the first time I have written this issue in such terms, but I am telling you it is inevitable. Why have I taken so long to discuss this issue? Well, over the years, I have brought up the topic of digital identity in humans and agents several times, largely out of the concerns I am sharing here. But indeed, I have fundamentally failed to approach it with the seriousness and urgency it requires. Frankly, I think many of my colleagues in the AI policy space have done the same. I believe this failure has two main causes.

First, these things are strange and off-putting, and many of us feel compelled to cater to the comfort zones of audiences rather than our own. Thus, among "serious people" (or those who want to be seen as serious), there is a universal tendency to confine candid discussions about the near future to private settings. Most of us know what the near future will look like. We relax in our small corners of Signal chats and Lighthaven, but when the public is watching, we discuss everything in more abstract, softer-sounding terms. This has been especially harmful in 2024 and 2025, when admitting any serious AI risks may be tagged as "apocalyptic" talk.

I am as guilty of this as my colleagues, if not more so. The problem is that being labeled a "crazy doomsayer" is unpleasant, and they say you wish to impose global fascism (and other similar, worse things). Constantly being labeled this way also limits a person's influence. Thus, many thinking about the governance of superintelligence, including myself, have avoided that unpleasantness, yielding to social pressures of self-censorship. I no longer do that, partly because I am tired of the straitjacket of discourse and partly because I now have an eight-month-old baby boy, and I must look him in the eyes every day.

Secondly, many believe that the arrival of what I call "self-sovereign AI" will constitute a catastrophic loss of control event. The worst-case scenario heralds the end of human existence, and the best scenario also heralds the end of human supremacy in the world. As my friend and former colleague Josh Achaim recently pointed out, a significant portion of the AI safety community holds these beliefs. They find it psychologically difficult to acknowledge an obvious truth: self-sovereign AI is coming, and it is coming soon. Let me be clear, I do not believe the end of human existence is likely to occur, but I fully acknowledge that the rise of self-sovereign AI could be extremely, extremely bad for humanity in some way.

I want to apologize for my failure to convey the specifics of self-sovereign AI with sufficient seriousness. I now understand this is a significant gap in my writing and speaking. Moving forward, I will strive to be more timely in noticing when I am biting my tongue, or worse, when I am closing my eyes.

免责声明:本文章仅代表作者个人观点,不代表本平台的立场和观点。本文章仅供信息分享,不构成对任何人的任何投资建议。用户与作者之间的任何争议,与本平台无关。如网页中刊载的文章或图片涉及侵权,请提供相关的权利证明和身份证明发送邮件到support@aicoin.com,本平台相关工作人员将会进行核查。

Share To
APP

X

Telegram

Facebook

Reddit

CopyLink