SpaceX Puts Starlink's Michael Nicolls Over xAI Data Centers After Reliability Failures
Elon Musk has installed senior Starlink executive Michael Nicolls to run SpaceXAI's data center operations after outages at the Macrohard facility repeatedly interrupted AI training runs and uptime fell below the 99.9% internal target.
SpaceX has carried out a significant leadership overhaul of the team building and operating its AI data centers, installing senior Starlink executive Michael Nicolls to oversee the xAI infrastructure division as reliability problems mount at its facilities in Tennessee and Mississippi, according to a report from The Information published September 1.
The reshuffle follows the departure of several senior data-center executives in recent weeks, and it lands at a moment when the central premise of xAI’s infrastructure strategy — that enormous AI compute clusters can be built faster and cheaper than the industry norm — is colliding with the unglamorous realities of operating them.
What happened
According to The Information, Elon Musk has shaken up the SpaceX team responsible for the company’s data centers in recent weeks, replacing several leaders after civil engineering concerns and consistent reliability issues surfaced at SpaceXAI’s existing sites. The company put Michael Nicolls, a longtime Starlink engineering executive, in charge of the AI data center organization.
Nicolls is a veteran operator. As VP of Starlink engineering he has publicly fronted the satellite network’s most difficult moments, including the July 2025 network outage that took the service down for roughly two and a half hours. Notably, he became xAI’s president in April 2026 and told staff in a memo that the near-term goal was to reach parity with Anthropic’s Claude models. Now the infrastructure side of the house answers to him too — a consolidation that mirrors how Musk historically runs programs: one senior lieutenant with end-to-end authority.
The timing is not accidental. The Information’s reporting describes facilities that were stood up so aggressively that some operated for months without backup cooling and power systems.
The Macrohard problem
The sharpest detail in the report concerns SpaceXAI’s Macrohard facility. To hit its extraordinary construction timelines, the site has leaned heavily on temporary infrastructure: mobile gas turbines for power, Tesla Megapack batteries for buffering, and more than 100 mobile chillers for cooling.
That improvised setup has consequences. Outages at the facility have repeatedly interrupted AI model training runs, according to a person involved with the project, and uptime has remained well below the company’s internal target of at least 99.9%. For a hyperscale training cluster, that gap is expensive. A cluster that goes dark mid-run doesn’t just pause — checkpointing and recovery can cost hours or days of GPU time, and at the scale of a gigawatt-class facility, every point of uptime translates into enormous sums of money.
The reliability issues have also coincided with delays to expansion beyond Memphis. But even against that friction, SpaceXAI says it had 1.4 gigawatts of capacity online as of June, is targeting more than 2 gigawatts by year-end, and generated roughly $2.6 billion in AI-related revenue.
The financial backdrop
The overhaul matters because AI has become one of SpaceX’s largest investment priorities almost overnight. In its most recently reported quarter, SpaceX posted $7.8 billion in revenue — nearly double the $4.1 billion from the same quarter a year earlier — while pouring about $15.83 billion into AI infrastructure during the quarter alone. The company remains loss-making even as its emerging AI revenue grows at roughly 250%.
Part of that revenue comes from eye-popping external compute deals. A $920 million-per-month agreement with Google gives the search giant access to roughly 110,000 Nvidia GPUs, and SpaceXAI has also leased Colossus capacity to Anthropic under a contract worth as much as $1.25 billion per month, though that deal includes cancellation provisions.
SpaceX has defended its build philosophy on cost grounds. The company said xAI brought the first two computing clusters at its second data center online for roughly one-quarter of typical industry costs — although its disclosure did not detail exactly which expenses were included in that calculation. If the number holds up under scrutiny, it’s a genuine structural advantage. If it omits the price of temporary power and cooling, redundancy, and the engineering rework now underway, the true economics are considerably less flattering.
Vertical integration, all the way down
Rather than slowing down, Musk is responding by pulling more of the data-center supply chain inside the company. SpaceX is developing manufacturing capacity for gas-turbine blades and vanes in Texas — an effort Musk said could shorten turbine deployment timelines by as much as 18 months in a market where power-generation equipment routinely faces multi-year waits.
And beyond Texas, the company is pursuing the most speculative leg of the strategy: orbital data centers. SpaceX has sought FCC permission for a system of as many as one million satellites intended to function as computing infrastructure in orbit. That filing, once a punchline, is now formally part of the same program whose ground-based facilities are struggling to hold a 99.9% uptime target.
Market reaction and what it means
Investors, for now, appear unbothered. SPCX traded around $143.36 on September 1, essentially flat against the prior close of $143.69. The shares sit roughly 36% below their June record of $225.64 but have recovered sharply from an August low near $104.83 — a range that suggests the market is pricing in both the upside of the AI build-out and the execution risk it keeps demonstrating.
The deeper significance is what this says about the industry’s speed-versus-reliability trade-off. SpaceXAI’s Memphis build-outs — Colossus, Colossus 2, Macrohard in Southaven, and a newly announced fourth Memphis site — were constructed at a pace the data-center industry considered impossible. The first Colossus supercomputer went from empty building to training cluster in 122 days. That speed let xAI secure scarce GPUs, lock in tenants like Google and Anthropic, and convert capacity into contracted revenue while rivals were still in permitting.
But the shakeup is an implicit acknowledgment that building fast and operating well are different disciplines. Hyperscalers like Google and Microsoft build slowly and redundantly precisely because a training cluster that flickers is a liability — their internal reliability standards have historically assumed years of commissioning work. SpaceXAI compressed that into months, and the bill is now coming due in the form of leadership churn, retrofitted cooling, and interrupted training runs.
The appointment of Nicolls signals that Musk’s response is not to moderate the pace but to professionalize the operations underneath it — bringing in an executive whose credibility was forged keeping a constellation of thousands of satellites online. Whether Starlink-style rapid iteration can be transplanted onto gigawatt-scale ground infrastructure is now one of the more consequential open questions in the AI build-out. The 2-gigawatt year-end target will be the first real test.
Sources
- [1] https://www.theinformation.com/articles/spacex-shakes-data-center-leadership-aggressive-build
- [2] https://coinpaper.com/35137/spacex-overhauls-ai-data-center-team-as-spcx-trades-near-143
- [3] https://finance.biggo.com/news/598184df-1a3e-4b4f-81e0-0f0a0d9f4bab
- [4] https://seekingalpha.com/news/4638950-spacex-is-said-to-overhaul-ai-infrastructure-team-as-data-center-delays-mount
- [5] https://www.kucoin.com/news/flash/spacex-overhauls-xai-leadership-amid-ai-infrastructure-challenges