Anthropic Slow AI Development Pledge
Anthropic slow AI development pledge to embed third-party evaluators could spur governance-focused investor scrutiny and tighter transparency demands.

KEY TAKEAWAYS
- CEO Dario Amodei pledged embedded permanent third-party safety evaluators and Anthropic committed to the first step.
- A researcher resignation and an internal estimate of more than 10.0% existential risk intensified governance scrutiny.
- The three-step framework lacks explicit speed limits or enforcement, leaving unclear measurement and escalation pathways.
HIGH POTENTIAL TRADES SENT DIRECTLY TO YOUR INBOX
Add your email to receive our free daily newsletter. No spam, unsubscribe anytime.
Anthropic’s slow AI development surfaced on Sept. 12 when CEO Dario Amodei published an essay calling for a deliberate slowdown and pledged to embed third-party safety evaluators. The statement followed a Sept. 9 researcher resignation and a public estimate assigning more than a 10% chance that advanced AI could pose an existential risk.
Amodei’s Pace the Frontier Plan
In his essay “We Must Pace the Frontier,” Amodei outlined a three-step framework to buy time for safety work through third-party oversight, industry coordination among democratic countries, and global government cooperation. He wrote, “We must slow the pace at which we improve the capabilities of AI models,” arguing that safety, alignment, and evaluation efforts are not keeping up with accelerating capability gains. The goal is to create breathing room for risk prevention.
Amodei distinguished pacing from halting progress, emphasizing that companies should take adequate time to align and safeguard models while allowing independent reviewers to confirm that work. The essay singled out recursive self-improvement—AI systems improving themselves—as a particular danger, warning such systems could advance drastically faster than anticipated and outrun human control. He sketched scenarios where novel cyberattacks, advanced biological threats, or autonomous “swarms” could compromise large parts of the internet within 6–12 months, causing damage measured in the hundreds of billions of dollars.
The first pillar of the plan calls for embedded third-party evaluators. Anthropic pledged to unilaterally adopt this model by hosting permanent external safety reviewers with employee-level access to systems, tools, offices, and development environments. These evaluators would verify adherence to safety measures, report incidents, assess alignment during training, and publish findings with minimal redaction and without Anthropic editorial control. Amodei named METR as a potential evaluator.
The second pillar urges coordination among leading AI firms in democratic countries to establish common safety standards and reduce commercial incentives to cut safety corners. Amodei noted some cooperation may require government support due to antitrust laws.
The third pillar envisions global coordination among governments—including outreach to countries such as China—to build international norms that constrain frontier AI development rather than leaving outcomes to whoever hosts the most compute. The framework does not specify numeric speed limits, capability ceilings, or enforcement mechanisms and acknowledges the measures vary in difficulty and need not be implemented sequentially.
Anthropic framed these commitments as operational safety measures rather than financial guidance, disclosing no revenue, margin, or earnings projections. The unilateral pledge to host embedded evaluators represents a direct operational change that could reshape transparency and governance norms at frontier AI labs.
Resignation and Risk Estimates
Jacob Coxon, a pretraining researcher who had worked at both Anthropic and OpenAI, announced his resignation and publicly warned that leading labs were racing toward self-improving superintelligence, describing the pace as a dangerous gamble. He accused developers of “gambling with our lives.”
Evan Hubinger, Anthropic’s Alignment Science Lead, publicly endorsed the seriousness of those warnings, stating the company “earnestly believe[s] AI could kill all humans.” He estimated more than a 10% chance that AI could cause human extinction within the next decade. Hubinger added that while current models pose limited risk, he worries about superintelligence arising from recursive self-improvement and said the company does not yet have a plan to solve alignment for superintelligence.
These statements have intensified public scrutiny of Anthropic’s governance, alignment strategy, and how the company balances frontier development with assessments of catastrophic risk.





