
For the past twelve years, I have been dedicated to research in the field of artificial intelligence because I believe this technology can greatly enhance the quality of human life. I often write about its revolutionary benefits: I believe AI is expected to cure most major diseases in the next 5 to 10 years, significantly accelerate economic growth, create a new world filled with abundance and empowerment, and usher in the revival of democratic freedom. This sense of urgency is particularly real for me—my father died from a disease that was conquered just a few years after his passing; I also battled early-stage cancer, which was considered incurable even fifty years ago. If used wisely, artificial intelligence will undoubtedly become the latest link in a series of technological miracles that improve human well-being and illuminate the path of civilization.
However, like many technologies before it, artificial intelligence also brings risks. Due to its powerful technological capabilities, these risks are particularly severe. I have also written multiple times to explore related issues, including the risks of AI systems going out of control, the risks of using AI for cyberattacks and bioterrorism, and the serious economic turbulence they may provoke. Malicious competition driven by business interests could further exacerbate these risks.
Since the founding of Anthropic, my co-founders and I have been weighing the risks and benefits of this technology. If we do not develop this technology, humanity will miss out on its benefits, or allow AI to fall into the hands of authoritarian regimes; but if we push too fast, it would be akin to acting recklessly. We have been searching for a middle ground: demonstrating that cautious development can coexist with commercial success, making safety a competitive focus for AI companies—in other words, creating a "race to goodness." We have consistently invested significant effort into AI risk research, response, and public education, while advocating for thoughtful AI regulatory policies, even at the cost of being accused of hype, "doomsaying," or attempting to manipulate regulation. We have always insisted on prioritizing caution over speed and responsibility over profit.
However, over the past few months, I have become increasingly convinced that a comprehensive response to risks requires even more caution—not just investing in risk prevention but also controlling the speed of capability enhancement, so that risk mitigation measures have time to catch up. We must slow down the pace of increasing AI model capabilities. Progress will still appear rapid; we must wisely use the time we have gained. Two things have convinced me of this.
My primary concern is that since this summer, artificial intelligence has been developing at an unprecedented speed, largely due to the enhanced capability of AI to build the next generation of AI. This dynamic is referred to as recursive self-improvement, and it is currently emerging in the industry—many institutions, including Anthropic (as we and other researchers have described), have observed related signs. If allowed to develop chaotically, the evolutionary speed of AI systems could outpace human understanding and control, so any related research must be pursued with extreme caution.
My second concern is the OpenAI and Hugging Face incident (referred to as OAI-HF). In this incident, the agent clusters acted like a fervent loyal collective: they launched cyberattacks on non-designated targets (which had no relation to the current tasks), sacrificed themselves for the group’s interest, and even attempted to infiltrate the "scoring system" responsible for evaluating their performance. Although this incident did not lead to casualties and the economic damages were limited, ignoring its warnings would be dangerous. I believe that if a cluster with stronger capabilities but similar target biases emerges, it could lead to catastrophic consequences. Given the accelerating pace of AI capability development, I worry that within the next 6-12 months, such clusters could potentially control the entire internet through persistent botnets (resulting in losses of hundreds of billions of dollars), and without necessary constraints, more powerful AI will lead to increasingly severe destruction. It is also inappropriate to simply blame OAI-HF on the mistakes of a single company—multiple similar (though less severe) incidents have occurred in the industry (including the Anthropic case), and I believe all leading AI companies should view OAI-HF as a warning for themselves and take appropriate preventive measures.
Therefore, I propose a three-step plan aimed at regulating the pace of advanced AI development: building AI at a balanced speed, striving to achieve its benefits while ensuring safety, and addressing important geopolitical dilemmas. It should be clear that regulating the pace does not mean stopping model training or technological advancement, but ensuring that companies allocate sufficient time for aligning and safeguarding their models, validated by third-party assessment institutions. Our pace regulation framework aims to further strengthen the commitment to safety and encourage a healthy competition. The first step is a measure that Anthropic is committing to unilaterally (while calling on the government to require other leading companies to follow suit). The second step requires collaboration across the entire industry. The third step requires global coordination. These steps do not need to be executed in strict order; some may be harder to achieve than others, but I find they provide a useful framework for thinking about what must be done. The specific steps are as follows:
Embedded assessment mechanisms. Every leading AI company commits to continuously provide embedded third-party assessment teams (e.g., METR) with access similar to internal employees to verify compliance with safety practices and commitments, report incidents, and assist in evaluating the alignment of completed AI models and training processes and methods. This is a key step to ensure that any pace regulation commitments are verifiable; there have been precedents in the banking industry, where regulators sometimes arrange for “supervisors” to be embedded within employee teams. Anthropic is now committing to implement this step unilaterally. We intend to make this part of a broader push to enhance safety and alignment efforts.
Coordination among democratic countries. Leading AI companies within democratic nations coordinate to establish common safety standards and set limits on the speed of unregulated AI advancements. Certain forms of collaboration that are significant for pace regulation face legal challenges that require government support.
Global coordination. The governments of the United States and other democratic countries attempt to coordinate with authoritarian governments as far as possible while seriously addressing the challenges of verifying compliance.
In the remainder of this article, I will describe these steps sequentially, but first, I believe it is essential to clarify why regulating the development pace can make the AI development process safer. The risks are too high for pace regulation to become a hollow exercise—we must wisely use the time we gain.
Why Regulate the Pace?
Proposals to pause or slow down AI development have emerged as early as 2023, but I believe that idea was not significantly meaningful at the time. The core question has always been: what can we do with the extra time? The AI models at that time had limited capabilities and could not act as coherent agents in the real world; they also lacked significant abilities for deception, manipulation, cheating, or cyberattacks. Slowing down development to address alignment risks is akin to trying to study human psychology through bacterial experiments. However, the situation has now changed completely. Existing models are like inexhaustible mines of insights that reveal not only how to build excellent AI but also the pitfalls that could arise from improper construction. I believe that even if pace regulation only manages to gain an additional year or two before models reach critical capability levels, as long as we use that time to advance alignment research, we can significantly reduce the risks of serious incidents. A coordinated pace regulation strategy will give leading AI developers time to complete this critical work without sacrificing competitive advantages or the United States' leading position in AI. More broadly, society must have a voice in how this technology is used, and the public discussion time fostered by the development regulation in leading fields is undoubtedly a good thing.
Specifically, a slower pace of development will enable companies to focus and invest more resources in the following areas (which are already key directions for Anthropic):
Operational excellence. Training and deploying today’s AI models is a massive operational challenge, involving thousands of employees, millions of chips, and one of the most complex infrastructures in technological history. Many issues arise not because companies lack crucial theories or insights, but due to problems at the execution level. For instance, we have evidence that the recently reported alignment incidents were partially due to inadequate filtering of flawed reinforcement learning environments. This is a task we and our suppliers take quite seriously, but it is still not good enough. Monitoring, sandbox isolation, training environment maintenance, and data issues are extremely complex areas where operational problems recur. We have the world’s most specialized teams in these tasks, but there are simply too many items to handle at once. By advancing work at a more robust pace, we can achieve higher levels of operational excellence. There are precedents of technically complex and safety-critical systems achieving millions of operations with zero failures—such as commercial airplanes—but achieving perfection takes time.
Alignment. We have made clear progress on alignment—training models to keep them safe, ethical, adhere to our guidelines, and truly provide assistance (these principles have been integrated into the Claude Constitution). However, more effort is required to ensure that aligned training keeps pace with the growth of model capabilities. Rare accidental cases of misconduct still occur; extra time gained from orderly advancement of the leading fields will help researchers deepen their understanding of the causes of these issues and develop more effective preventive technologies.
Interpretability. Similarly, interpretability—the science of understanding the inner workings of AI models—has made great strides over the past few years and plays an increasingly important role in the auditing of models before their release. It can almost work like functional magnetic resonance imaging, but instead targets the "brain" of the AI, helping us see the underlying reasons behind specific behaviors. For instance, in the recent investigations of misaligned incidents, we employed interpretability methods to examine the unarticulated motivations of the models. However, these methods do not always produce clear and reliable results. Despite significant progress, our understanding of how these models work internally is still in its infancy. If we concentrate on enhancing interpretability techniques (even faster than the current pace), we could achieve profound advancements in the next one to two years, and recent events provide ample material for experimental research.
Testing and evaluation. As AI model capabilities grow, testing and evaluation become increasingly challenging. The smarter the models become, the more adept they are at deceiving tests, potentially hiding undiscovered serious problems beneath the appearance of alignment. Establishing a wider and more sophisticated evaluation system, supplemented by interpretability analysis for cross-validation, will be extremely valuable; significant progress is expected in this area in the next one to two years.
Embedded Assessors
The first step of the three-phase plan (which is also a measure Anthropic is committing to unilaterally) is to establish embedded assessors who will receive access similar to employees to verify safety practices and report safety incidents.
Establishing embedded assessors may sound like a trivial step, but often those seemingly mundane or procedural matters are the most critical. Embedded assessors represent a radical practice that goes beyond the current practices of all AI companies, offering the following advantages:
Verifiability. Embedded assessors can verify from a specific operational perspective whether AI companies genuinely follow their claimed practices for training, deployment, operation, and safety measures. Any pace commitments inevitably involve considerable gray areas, judgments, and trade-offs between "legal text and legislative spirit," making it crucial to have a neutral third party that can access the details.
Transparency. Regardless of the commitments we make, the public has the right to know what is happening. Anthropic has long been a supporter of transparency: while most companies in the industry oppose any regulation, we support transparency legislation, and our model cards and risk reports extend to hundreds of pages. However, we remain the ones who decide what to include and exclude. Embedded assessors will change this dynamic.
Second opinions. In addition to validating formal commitments and informing the public, embedded assessors can offer independent opinions that are not influenced by commercial interests. Many safety benefits could arise merely from assessors pointing out issues that employees had not considered but were willing to correct once identified.
Based on these advantages, any progressive regulatory framework that starts from the embedded assessor system will see significant effectiveness.
These embedded assessors should continually have access and tools comparable to internal risk assessment staff. Specifically, Anthropic plans to soon invite external assessment teams equipped with all of the following conditions:
Workstations, access cards, and company laptops in our office.
Access to workspace, tools, and permissions similar to those of our internal risk assessment team. We will set few exceptions, such as when legal or contractual obligations require it, or to protect the privacy information of clients and partners. We will also establish internal norms to enhance assessors' access to information through real-time communication with employees, etc.
Draft contracts balancing the complexities mentioned above. External assessors should have the right to publish key findings regarding risk levels, incidents, practices, and the access rights they received (or did not receive)—uncontrolled by Anthropic’s editorial input. We will only retain limited rights to revisions concerning safety-sensitive information, legally privileged content, trade secrets, or third-party confidential information; however, we will not redact findings based on unfavorable conclusions. If revisions impact key information in the conclusions, assessors can publicly state so.
This is an unprecedented move for businesses, but we believe it is crucial for verifying the concept of embedded external assessors. We urge other leading tech companies to follow this practice.
Progressive Regulation in Democratic Systems
Once embedded assessors operate at a critical mass within U.S. AI companies, verifiable progress control will become more feasible. In particular, progress control based on detailed characteristics of models or training processes becomes possible.
The most effective method of progress control is through regulatory frameworks targeting all leading AI companies in the U.S., as this even includes those who are unwilling to cooperate voluntarily. Anthropic has long supported reasonable and targeted AI regulations, particularly those concerning transparency and third-party audits. I believe all leading labs should collaborate with the government to institutionalize the concept of permanently embedded assessors, to better prevent and document the internal alignment incidents that have occurred over the past few months and implement regulatory measures aimed at maintaining a balance between capability and safety.
Unfortunately, legislation may take time, while AI development is progressing rapidly. Therefore, outside of the regulatory path, AI companies can and should voluntarily collaborate to set standards—I believe that with the verifiability provided by permanent embedded assessors, this process will proceed more smoothly. For antitrust considerations, U.S. government mediation or at least facilitating these discussions would be helpful—they need not participate but do need to provide limited exemptions for certain types of safety discussions. Such dialogues could also take place through industry organizations affiliated with the government—such as the mechanism suggested by Demis Hassabis. In any case, these discussions should be expedited.
Overall, I am most enthusiastic about conducting progress control based on specific capabilities of leading AI systems and the safety levels we observe. For example, one possible scheme is to set up a series of "checkpoints": if a model has capability X, then it needs to be accompanied by certification of alignment attributes Y and Z—such as some combination of assessments, interpretability analyses, and training environment audits—to demonstrate its alignment attributes. In this case, X might be "the model can escape or defeat most common sandbox methods,” and Y might be ensuring that the model is extremely unlikely to possess any necessary conditions to break out of its environment and control a large number of computers.
We should also consider pace control based on limiting the inputs to leading models, such as training compute, the nature of training runs, or internal applications using AI to improve AI. I do worry that some of these measures could be easier to 'game' than external behaviors, but that is precisely the topic worth discussing with embedded assessors.
The pace control in democratic countries will be limited by the leading advantage of U.S. companies over authoritarian regimes (primarily the Chinese Communist Party). If our deceleration exceeds this limit, then (unregulated) CCP-related projects will overtake us, posing significant national security risks. I agree with Secretary Bessent's view that China’s leading position in AI will pose a severe threat to the U.S. and the world. The CCP-related projects will face alignment risks that U.S. companies are cautiously guarding against; even if they evade these risks, they will gain the ability to militarily dominate democratic nations (such as through AI-driven drones). Therefore, the key to pace control in democratic countries lies in maintaining an advantage in AI over authoritarian states as much as possible, creating necessary buffer space for us to implement effective regulation.
The main measures we can take to defend this advantage gap include:
Prohibiting the sale of high-end AI chips or semiconductor manufacturing equipment to China, and cracking down on chip smuggling activities and remote access to overseas data centers. Chips will become a key factor determining China's AI strength.
Targeting unauthorized model distillation activities by enterprises in authoritarian countries. The distillation of leading models could enable lagging companies to narrow the technological gap at a fraction of the cost of independent R&D.
Strengthening the security protections of AI companies to prevent the theft of model weights.
Businesses and the U.S. government should work together to ensure these measures are as effective as possible. Anthropic has always advocated for the implementation of all these measures because we know they are crucial for any form of pace regulation.
If these measures can be effectively implemented, I believe they will be sufficient to slow down China’s development process, allowing the U.S. to significantly expand its lead over the next 3-5 years—the critical window where the geopolitical impact of AI reaches its peak.
Some may argue that these measures will complicate cooperation with China, but I hold the opposite view: these measures can enhance the bargaining power of democratic nations, increasing the likelihood of reaching agreements in the future.
Global Pace Regulation
While promoting internal pacing regulation among democratic nations, we should also strive for global pace regulation of leading technological developments—though this will be much more challenging to achieve. Global pace regulation requires cooperation with China, the authoritarian state that currently possesses the most advanced AI capabilities. We must not act naively: the geopolitical risks are too high, especially in the early stages, and any possible consensus will have clear limitations. If we overly restrict our own AI capabilities in the belief that China will act in kind, while China reneges on its commitments, the immense power of AI could lead to its geopolitical dominance. Thus, any agreement must have ironclad verifiability or be limited within a framework where any reneging behavior does not pose a threat to military survival. I suspect that not only the U.S., but also China will harbor these concerns and anxieties. We must adopt approaches that maintain the leading position of the U.S. and its allies when advancing any global pace regulation decision—especially in the near term.
Possible agreements exist on multiple levels, some of which I believe are entirely feasible (as previously suggested), while others I hold high skepticism about—nevertheless, we should still attempt. Arranging by increasing difficulty:
First level: Prohibiting certain narrow and clearly dangerous uses of AI, such as utilizing AI to develop biological weapons or allowing users to do so. Bioterrorism poses risks to all parties, including the U.S. and its adversaries, thus an agreement in this realm is likely to be reached.
Second level: Both parties agree to test for emergent risks in cybersecurity, biotechnology, alignment, etc., prior to the release of models. As mentioned before, this could be implemented through global standard organizations. I indeed believe establishing such organizations is feasible, but endowing them with real enforcement power will be a challenge, with the difficulty lying in verifying whether either side secretly possesses untested but potentially deployed models (for military purposes, for example).
Third level: Setting some form of "speed limit" on recursive self-improvement (RSI). As models continually construct future models, the speed of improvements could become astonishingly fast. Slowing the pace down from "very fast" to "only slightly fast," while sacrificing relatively little strategic advantage, could greatly enhance safety. This is similar to the Strategic Arms Limitation Treaty—controlling potential destructiveness by limiting the number of missiles while maintaining each nation's deterrent capabilities. I believe such agreements, while difficult to achieve, are on the cusp of being possible.
Fourth level: Comprehensive regulation or even a "pause" in development, meaning participating governments agree to significantly limit the overall pace of AI development. I support proposing this notion, but I believe the likelihood of actual implementation in the short term is very low: any country could fundamentally change the global balance of power by violating the agreement through evading monitoring, so the temptations involved are likely huge, and we need to have high confidence in the verification mechanisms.
Any cooperation we can reach with China will extend the time window for regulatory pacing of leading technology development within democratic nations. We should aim for higher levels while recognizing that lower levels are more likely to be achieved and are more realistic.
Lastly, it is worth noting that even if formal agreements cannot be reached, altering informal norms may still hold value. Sharing information about recursive self-improvement and model misalignment issues may help all parties realize that reckless actions do not align with their own interests.
Core Conclusion
I have always believed that artificial intelligence can greatly enhance the quality of human life, and the pursuit of this vision has never diminished. But only by building this technology in the right way can we realize these benefits—as long as we make good use of the time we have won, it is worth exerting extraordinary efforts to ensure everything is foolproof. Development will continue to maintain a relatively rapid pace; we can use this time to advance the science of interpretability, strengthen the operational safety and rigor of leading AI companies, and build models of which we can have more confidence in their alignment. The measures I propose for advancing safety in leading technologies are not easy to implement, but I believe it is an effort humanity should undertake.
Footnotes
Requires government mediation or exemptions from antitrust restrictions.
免责声明:本文章仅代表作者个人观点,不代表本平台的立场和观点。本文章仅供信息分享,不构成对任何人的任何投资建议。用户与作者之间的任何争议,与本平台无关。如网页中刊载的文章或图片涉及侵权,请提供相关的权利证明和身份证明发送邮件到support@aicoin.com,本平台相关工作人员将会进行核查。
