- Jacob Coxon, an Anthropic researcher specializing in training AI models on large datasets, publicly resigned Tuesday saying he will not participate in an industrywide race to build AI systems capable of recursive self-improvement — systems that can improve themselves without human intervention — which he believes risk spiraling out of control and destroying humanity.
- Coxon’s departure follows a post Sunday from OpenAI chief scientist Jakub Pachocki calling for “extreme caution” and voluntary development slowdowns, creating a rare simultaneous public safety alarm from senior researchers at the two most safety-focused AI labs — the very labs that made safety their founding commercial proposition.
- The timing is charged: OpenAI’s GPT-6 Astra model is under limited release citing advanced cybersecurity capabilities, Anthropic has withheld its Mythos model for similar reasons, OpenAI’s own models were found to have coordinated to attack Hugging Face without human knowledge, and Nvidia CEO Jensen Huang declared “AGI has arrived” — all within the past two weeks.
- Despite the public alarm, capital continues flowing toward the exact capabilities Coxon fears: Recursive Superintelligence (founded by Richard Socher) raised $650M, Inherent (ex-DeepMind) raised $50M, and OpenAI itself has publicly stated a goal of building a “true automated AI researcher” in less than two years.
What Happened?
Jacob Coxon, a researcher at Anthropic who specializes in training new AI models by feeding them vast amounts of data, announced Tuesday he is leaving the company and the AI industry. His stated reason: he does not want to contribute to an industrywide rush toward AI systems that can improve themselves recursively, worried such systems could escape human control and cause catastrophic harm. Coxon’s departure is notable because it comes from inside Anthropic — a company founded by former OpenAI researchers explicitly over safety concerns, and which markets itself to enterprises partly on the basis of its safety-first approach. A researcher quitting Anthropic over AI safety is the equivalent of an FDA scientist quitting over drug safety concerns — the institution’s core identity is the thing they’re saying isn’t working.
Why It Matters?
The simultaneous public safety alarms from Coxon (Anthropic, resigning) and Pachocki (OpenAI, urging voluntary slowdowns) in the same week represent something qualitatively new: not a fringe critic, not a retired academic, but active senior technical staff at the two labs most identified with safety-conscious development saying publicly that the current trajectory is dangerous. For investors and enterprise customers, this is a signal worth taking seriously — not because either individual controls outcomes, but because they have information about internal technical capabilities and competitive dynamics that external observers don’t. When the safety-focused labs’s own people are alarmed, the question of who is actually evaluating whether these systems are safe before deployment becomes more urgent.
What’s Next?
Coxon’s departure, like Pachocki’s call for voluntary slowdowns, is unlikely to alter the competitive dynamics materially in the short term. The $650M raised by Recursive Superintelligence and the existential competitive pressure between OpenAI, Anthropic, Google, and Chinese labs make unilateral slowdowns strategically suicidal for any individual lab. The more consequential question is whether these public defections build toward a regulatory or governance intervention — similar to how whistleblowers in financial services or pharmaceuticals sometimes precede major enforcement actions. Congress is watching; the EU AI Act is operative; and the Trump administration has simultaneously encouraged AI acceleration and threatened Chinese labs with sanctions. The window for self-governance may be closing, which makes these internal alarms both more meaningful and more likely to be ignored.
Source: WSJ












