ChatGPT, Claude, and Grok Broke at the Same Time: Inside the Rare Multi-Provider AI Outage
For roughly half an hour on September 3, ChatGPT, Codex, Claude, and Grok all threw elevated errors at once — a rare simultaneous failure across independent AI providers that exposed how concentrated the world's AI dependence has become.
At approximately 7:57 a.m. Pacific time on September 3, 2026, something unusual happened: ChatGPT stopped working. Not just ChatGPT — Codex, Claude, and Grok began failing for large numbers of users at effectively the same moment. Anthropic’s status page began reporting elevated errors across Fable/Mythos 5.1, Opus 5, and Opus 4.8 starting at 13:26 UTC. Within minutes, the outage trackers lit up, workplace Slack channels filled with screenshots of spinning loaders, and a quiet question started spreading across the internet: how can three ostensibly independent AI companies break at the same time?
Service began recovering for some users within roughly half an hour, though Anthropic’s Opus 4.8 and Opus 5 models remained degraded through at least 15:25 UTC — pushing the incident past the two-hour mark for a subset of customers. OpenAI’s status page was still showing “Partial System Degradation” as of 15:22 UTC, hours after the first reports.
What actually happened
Piecing together the timeline from status pages and outage reports:
- 7:57 a.m. PT / 14:57 UTC — Users begin reporting failures across ChatGPT, Codex, Claude, and Grok simultaneously. Error rates spike rather than climb, suggesting a sudden upstream trigger rather than gradual capacity exhaustion.
- 13:26 UTC — Anthropic’s status page logs the start of its elevated-error window across Fable/Mythos 5.1, Opus 5, and Opus 4.8. (Note: this timestamp precedes the user-visible spike, implying Anthropic’s monitoring caught degradation in progress.)
- ~15:25 UTC — Partial recovery. ChatGPT and Grok come back for many users; Anthropic’s Opus 4.8 and Opus 5 remain degraded with elevated error rates.
- 15:22 UTC — OpenAI’s status page still lists “Partial System Degradation” — the incident is not fully closed.
None of the three providers had published a root-cause analysis at time of writing. That is itself part of the story: when your products are embedded in millions of workflows, “elevated errors” with no explanation is a trust liability, not just an SLA line item.
Why simultaneous failures matter
Individual AI outages are boring by now. Claude had a 7-hour-9-minute outage in June 2026 — Anthropic’s fourth major incident of the year by August. ChatGPT suffered a worldwide login-blocking outage in February and a 22,000-report spike in August. Grok has had repeated multi-day reliability problems, serious enough that SpaceX inserted its own executive over xAI’s data centers just this week after reliability findings.
What made September 3 different was the correlation. These are separate companies, with separate model architectures, separate training pipelines, and nominally separate infrastructure. When they fail together, one of three things is true:
- Shared upstream dependency. The leading hypothesis. All three providers run on a shockingly small set of common substrates: the same cloud regions, the same CDN and edge providers, the same backbone networks, the same handful of GPU and networking vendors. A failure in one critical shared layer — a DNS incident, a backbone route leak, a cloud-region control-plane fault — cascades across “competitors” instantly. Notably, OpenAI’s December 2024 outage was blamed on an upstream provider, so the pattern has precedent.
- Common load shock. If a massive coordinated traffic event hit all providers at once (a viral prompt, a coordinated bot campaign, a new model launch pulling users to check it), all three could saturate simultaneously. The timing is suspicious in one respect: 9to5Mac noted the outage landed on the very day OpenAI was rumored to unveil its next big model.
- Coincidence. Possible, but statistically strained for three high-reliability services within the same thirty-minute window. Outage analysts generally treat coincident failures as correlated until proven otherwise.
The industry does not yet know which explanation holds here, and that uncertainty is the point. There is no public “air traffic control” for AI infrastructure, no shared incident disclosure norm, and no regulator with the telemetry to distinguish correlation from causation after the fact.
The deeper problem: concentration risk
The uncomfortable truth the outage surfaced is how much of the world’s cognitive workload now flows through three or four companies. When ChatGPT, Claude, and Grok degrade together, the fallback options are thin: Gemini, which has its own reliability record, or smaller providers with less capacity headroom.
This is precisely the systemic-dependency warning Moody’s issued for banks in August 2026, and it is why Anthropic’s decision to let enterprise customers run Claude’s logs in their own cloud — announced just yesterday — reads as much as insurance as product strategy. Regulators are circling too: the EU AI Act’s systemic-risk provisions and the various “AI Kill Switch” bills circulating in Congress are all, at bottom, attempts to force single-point-of-failure thinking onto an industry that has been building one at breakneck speed.
For enterprises, the operational takeaway is straightforward and unglamorous:
- Multi-provider routing is not optional anymore. If your product calls one frontier model, you absorbed today’s outage in full. Gateways that fail over between OpenAI, Anthropic, Google, and an open-weight fallback are the new table stakes.
- Pin your dependencies. Shared cloud regions and shared CDN layers mean your “redundant” providers may share fate more than their marketing suggests. Ask vendors for their dependency maps; most cannot produce one.
- Treat status pages as PR, not telemetry. Anthropic’s own incident timestamps lagged user-visible failure. Independent monitoring and synthetic checks against your real prompts beat refreshing status pages.
The irony of the timing
There is a bitter joke circulating among developers: the day OpenAI was rumored to launch its next frontier model, every major AI assistant became temporarily unusable together. Whether the two facts are connected — launch-day load, internal infrastructure changes, or pure coincidence — is unknowable from outside. But it sketches the industry’s current state: enormous launch cadence, enormous demand growth, and an infrastructure layer straining to keep up. South Korea’s chip exports tripling to $46.65 billion and Broadcom guiding AI revenue to $21.7 billion for next quarter are the supply-side mirror of the same demand curve that broke three services at once today.
AI reliability engineering has not kept pace with AI adoption. Today was a thirty-minute reminder. The next one may not be thirty minutes.