- Rapid swings in AI data centers’ power demands are physically destroying the equipment inside them: gas turbine cranks have broken off at multiple facilities; gas-fired turbines at xAI’s Colossus computing facility in Memphis have developed cracks; batteries installed to smooth power fluctuations are failing within weeks or months rather than their designed lifespans; cooling systems are wearing out far ahead of schedule; the underlying cause is the extraordinary volatility of AI training workloads — when training new models, hundreds of thousands of GPUs power up and down on a millisecond basis, causing power usage to spike as much as 50% above design capacity in an instant; Drew Baglino, former Tesla executive now at Heron Power Electronics, describes a 1 GW facility using 1.5 GW “for a split second” — equipment built for steady-state loads is being subjected to repeated shock cycles it was never designed to absorb.
- The scale of the power volatility makes the analogy to conventional industrial equipment apt but insufficient: a 1 GW AI data center is equivalent in power consumption to the entire city of Boston, and some planned Texas and Midwest AI campuses will consume nearly as much power as New York City on average; the training process causes what Shannon Miller of Mainspring Energy describes as “half of Boston flickering on and off every few seconds”; the power management challenge is not just about raw capacity but about rate of change — Jon Parrella of Terraflow Energy compares it to shifting a Ferrari directly from sixth gear to first: “You can’t swing that fast”; the stabilizing equipment (batteries, capacitors, flywheels, transformers) that exists to buffer these swings is being deployed insufficiently at new data centers being built as fast as possible, and even where it is installed, the extreme duty cycle is burning it out faster than operators anticipated.
- The reliability problem is already hitting revenue: data centers are financed and contracted on the assumption of 99.999% uptime — essentially continuous operation 365 days per year; in reality, some AI computing facilities are seeing uptime closer to 80%, according to a person involved in data center financing; the financial consequence is not the replacement cost of failed equipment but the value of expensive GPU compute capacity that is offline and not generating revenue; downtime costs range from thousands to hundreds of thousands of dollars per minute depending on the facility and workload; Microsoft’s planned 2.67 GW AI campus in West Texas has already been delayed from 2027 to 2028 specifically because additional engineering time was needed to achieve the 99.999% reliability requirement — a delay that itself represents a revenue timing hit for the developer.
- The grid stability dimension is potentially the most consequential and least-appreciated risk: AI data centers are large enough that their power swings don’t just stress their own internal equipment — they create sub-synchronous oscillations in power flows that can damage equipment connected to other parts of the network beyond the data center fence; NERC, the top US grid reliability regulator, found that approximately three-quarters of operational data center load models “are insufficient to represent data-center dynamic behavior” and issued a rare level-3 alert requiring big data centers to address these risks immediately (response deadline: August 3); the combination of AI data centers causing internal equipment failure, revenue-destroying downtime, and external grid instability suggests the total cost of the AI infrastructure buildout is being systematically underestimated by hyperscalers, investors, and lenders alike.
What Happened?
AI training workloads are spiking data center power use up to 50% above design capacity within milliseconds — cracking gas turbines (including at xAI’s Memphis Colossus), burning through batteries in weeks, and wearing out cooling systems prematurely. Some facilities are delivering ~80% uptime versus the 99.999% promised. The power swings are large enough to cause sub-synchronous grid oscillations that damage external equipment. NERC issued a rare level-3 alert requiring immediate remediation, finding ~75% of data center load models inadequate. Microsoft’s West Texas 2.67 GW campus has been pushed back a full year.
Why It Matters?
The AI infrastructure investment thesis depends on data centers operating continuously to generate returns on hundreds of billions in GPU capex. Eighty-percent uptime versus 99.999% contracted uptime is not a rounding error — it’s a fundamental threat to the economics. The equipment failure problem also compounds over time: as turbines crack and batteries fail prematurely, maintenance costs rise, replacement cycles shorten, and the depreciation schedules that underpinned lender underwriting become inaccurate. Add grid stability risk and you have a sector where the physical infrastructure is systematically underperforming the financial models built around it.
What’s Next?
Watch NERC’s August 3 deadline response — which facilities submitted plans and what remediation they’re committing to; watch hyperscaler earnings commentary for any acknowledgment of uptime shortfalls or unplanned maintenance costs; watch insurance rates for AI data center coverage as a market signal of how underwriters are pricing the equipment failure risk; watch Nvidia’s power delivery work on its 2027 next-generation servers (which Heron Power Electronics is developing equipment for) to see whether the industry is building solutions fast enough to stay ahead of the problem; and watch whether the 80% uptime reality begins surfacing in lender covenant violations or developer-hyperscaler contract disputes.
Source: Bloomberg













