- Evan Hubinger, an alignment lead at Anthropic, said publicly that he places the odds of AI killing all humans within the next decade above 10%. The comment came after a colleague resigned over safety concerns and triggered a wave of similar warnings from researchers at both Anthropic and OpenAI.
- The specific worry is recursive self-improvement, where AI systems help build their own more capable successors. Anthropic said in June that its internal data shows Claude accelerating AI development and that the trend is moving faster than the company expected. Its engineers now ship roughly eight times as much code per quarter as they did between 2021 and 2025.
- OpenAI’s chief scientist Jakub Pachocki wrote that nobody is prepared for the consequences of continued rapid gains in machine intelligence, and that systems arriving in the next few years will increasingly drive their own development. Researchers at both labs said publicly that no viable scientific plan yet exists for managing these risks.
- The warnings land against record AI capital deployment. TSMC’s August revenue rose more than 53% to a record on chip demand, Google committed at least $15 billion to AI infrastructure in Finland, Mistral reached a $24 billion valuation on a Samsung-led round, and Qualcomm issued Amazon warrants for $4 billion of stock as part of an infrastructure deal.
What Happened?
Safety warnings from inside the two leading AI labs went viral this week following the resignation of an OpenAI researcher over safety concerns. Anthropic’s Hubinger clarified that his concern centers on superintelligence emerging from recursive self-improvement rather than current systems. Researchers at both companies amplified the point on social media. Anthropic published a blog post in August laying out three scenarios: frontier progress stalling with capabilities widely diffused, which it considers unlikely; continued gains with humans in control, which it calls likely; and full recursive self-improvement with humans playing a substantially diminished role in development, where it says it is least certain how the alignment problem resolves.
Why It Matters?
The notable feature here is the source. These are not external critics but the people building the systems, and both companies are on record saying autonomous improvement is arriving ahead of their own forecasts. Carnegie Mellon’s Vincent Conitzer noted that AI is already contributing genuinely new ideas, which makes the point of acceleration very hard to forecast. For investors, this creates an unusual tension: the same capability curve driving TSMC’s record revenue and hyperscaler capex is the one prompting the warnings. That is a regulatory risk vector that does not appear in any current model. Cohere’s Aidan Gomez separately described leading AI models as among the most potent cyber weapons ever built, and said the US lead over Chinese labs is evaporating quickly — which complicates any policy response built on unilateral restraint.
What’s Next?
Watch whether these warnings translate into policy movement or remain confined to research discourse, since the labs are effectively flagging risk in their own product. Further departures from safety teams at either company would be the clearest signal that internal disagreement is escalating. On the competitive side, the sovereign AI positioning that Cohere and Mistral are pursuing gains relevance if data governance concerns harden, and Samsung’s backing of Mistral suggests capital is already moving that way. Anthropic also disclosed that Chinese labs used millions of Claude exchanges to train their own models, which points to distillation and model security becoming the next enforcement battleground.
Source: CNBC














