Anthropic Resignation Amplifies Safety Concerns
Anthropic Resignation spotlights internal safety warnings and a public alignment lead assessment, raising regulatory and investor scrutiny.

KEY TAKEAWAYS
- Anthropic resignation highlighted a pretraining researcher warning labs are racing toward self-improving superintelligence.
- Alignment lead Evan Hubinger posted a personal greater-than-10% extinction-risk estimate for AI within the next decade.
- Hubinger said Anthropic lacks a clear plan to solve alignment for future superintelligence, increasing scrutiny.
HIGH POTENTIAL TRADES SENT DIRECTLY TO YOUR INBOX
Add your email to receive our free daily newsletter. No spam, unsubscribe anytime.
Anthropic’s resignation on Sept. 8 involved a pretraining researcher stepping down and a widely viewed post by the firm’s alignment lead warning of significant long-term risk from advanced AI, intensifying scrutiny of the company’s safety trajectory and potential regulatory exposure.
Resignation Highlights Internal Alarm Over AI Race
Jacob Coxon, who spent roughly three years on pretraining research at OpenAI and Anthropic, announced his resignation on Sept. 8. He said he no longer wanted to accelerate model-training work he judged dangerous and irresponsible. Coxon warned that leading labs are racing toward self-improving superintelligence—AI systems that can recursively enhance their own capabilities—and argued that safety efforts have not kept pace with these advances. He described the situation as an unacceptable “race” and urged broader coordination among developers. Coxon called for private companies to avoid unilaterally initiating an “endgame” for superintelligent systems, warning that such a move is a “hubristic gamble.” No formal Anthropic corporate response to Coxon’s resignation or the public posts by researchers has been identified.
Alignment Lead’s Public Extinction-Risk Estimate
Evan Hubinger, who leads Alignment Science at Anthropic, posted on Sept. 8 at approximately 7:27 p.m. ET that he personally assesses a greater-than-10% chance that AI could kill all humans within the next decade. He emphasized this figure reflects his own view rather than an official company forecast. Hubinger said he considers the risk from current consumer AI models to be low but is specifically concerned about future superintelligence arising from recursive self-improvement. He added that Anthropic does not yet have a clear plan to solve alignment—the process of ensuring AI systems act in accordance with human values—for such advanced systems. Hubinger’s post brought the risk of AI-driven extinction into broader public view. No new regulatory filings, enforcement actions, or formal policy changes tied directly to the resignation or Hubinger’s estimate have been identified.
Hubinger stated, “We really do earnestly believe AI could kill all humans.”





