Two Chips or Four: Apple Weighs a 2029 Return to the Server Market with M8 Ultra and NVIDIA's NVLink
Apple is reportedly designing enterprise AI servers with two or four M8 Ultra chips linked by NVIDIA's NVLink Fusion — its first server product since Xserve, aimed squarely at AI inference and EU data sovereignty.
For a company that has spent a decade telling the world it doesn’t need the data center, Apple is suddenly spending a lot of time thinking about racks. According to a report from The Information published September 16, 2026 — and confirmed in outline by Bloomberg — Apple has been working on plans for an enterprise server built on its own silicon, and has held discussions with NVIDIA about using its networking technology to tie the machine together. If it ships, it would be Apple’s first commercial server since the Xserve was discontinued in 2011, and one of the strangest bedfellow stories of the AI boom: the company that famously refused to buy NVIDIA GPUs, quietly licensing NVIDIA interconnect.
What the report actually says
The details, sourced to people familiar with the plans, sketch a machine with a familiar Apple shape: minimal configurations, maximal polish. Two variants are under consideration — one pairing two M8 Ultra chips, and a larger version pairing four. The servers would target the AI inference market, aimed at AI developers, businesses, and potentially regulated sectors that need to run large models on-premises. NVIDIA’s NVLink Fusion technology would be evaluated to connect the M8 chips into a coherent whole, letting multiple SoCs behave like a single, larger accelerator with shared memory bandwidth.
The timeline is the sobering part: nothing is expected to reach the market before 2029, and The Information notes the project could still be canceled before it ever materializes. Apple has a long history of exploring server-class hardware and shelving it — the M2 Ultra-based server racks Apple built for its own Private Cloud Compute were reportedly judged underpowered for frontier-scale AI workloads, one of several factors that pushed the company toward Google TPU capacity and, now, toward selling inference boxes instead of only renting them.
Why M8 Ultra, and why NVLink
The choice of silicon is the least surprising part. Apple’s Ultra-tier SoCs — two Max dies fused via a silicon interposer — have always been the closest thing consumer silicon gets to data-center parts: enormous unified memory pools, exceptional performance-per-watt, and no dependence on HBM supply chains that are booked out years in advance. An M8 Ultra with a large unified memory address space is a credible inference engine for large language models, particularly at a time when memory capacity, not raw FLOPS, is the binding constraint on serving big models economically.
The NVIDIA angle is where it gets interesting. NVLink Fusion is NVIDIA’s semi-custom program — opened in 2025 — that lets third-party CPUs and accelerators plug into the NVLink fabric that normally binds NVIDIA’s own systems together. For Apple, the attraction is obvious: building a coherent multi-chip fabric is brutally hard, and NVIDIA has already solved scaling, switching, and the software stack that comes with it. For NVIDIA, licensing interconnect to a company that wouldn’t buy its GPUs is a pure-margin win and a signal of how the “AI factory” market is segmenting — you no longer need to sell the whole rack to participate in the rack.
It is also a striking reversal. Apple and NVIDIA have been estranged for nearly a decade, with Apple declining to adopt NVIDIA GPUs in Macs since 2018 and the two trading barbs over driver support and market power. A 2029 server jointly architected around Apple compute and NVIDIA fabric would be the most significant technical collaboration between the two companies in modern memory.
The sovereignty angle
Reporting around the project suggests European data rules are part of the motivation. The Next Web’s coverage notes the M8 Ultra server could be positioned for EU sovereignty requirements — on-premises boxes that keep regulated data entirely within a customer’s physical control, something cloud GPU rental struggles to guarantee. Apple’s privacy-first brand, built around Private Cloud Compute and on-device processing, maps naturally onto “the AI server you can actually trust” as a pitch to banks, hospitals, and governments.
That positioning would also differentiate Apple from the incumbent inference players. NVIDIA sells accelerators; hyperscalers sell rented capacity. Apple would be selling something closer to what it has always sold: an integrated appliance, hardware and software co-designed, with the model runtime treated as part of the product rather than a compatibility burden.
Context: Apple’s belated infrastructure pivot
The report lands mid-pivot for Apple. After years of being perceived as behind on generative AI infrastructure, 2026 has seen the company lock in Google TPU capacity for foundation-model training, continue work on its dedicated “Baltra” AI server chip with Broadcom (itself reportedly delayed, with acquisition of chip startups floated as a fallback), and ship the next generation of Apple Intelligence with an on-device-forward architecture. Selling enterprise inference servers would be the logical next step: if Apple has to build AI infrastructure for itself anyway, the marginal cost of productizing it is falling.
Skeptics will note the obvious risk: 2029 is a geological era in AI time. Whatever the inference market looks like when the machine ships, it will not look like today’s. The two-vs-four-chip configurations read as conservative engineering for a market that is currently rewarding 72-GPU racks and gigawatt clusters. Apple is betting that by decade’s end, a meaningful slice of AI compute will have moved back on-premises — pushed by cost, privacy law, and simple distrust of vendor lock-in — and that performance-per-watt and integration will matter more than peak throughput.
What to watch
Three signals will tell us whether this is real: further NVIDIA Fusion design wins beyond the CPU partners already announced; whether Baltra’s timeline firms up or collapses into the M-series server line; and any Apple enterprise-sales hires or certifications (FedRAMP-equivalent, EU data residency programs) that would precede selling racks to strangers. Until then, it remains one of the more intriguing “reported plans” of the year — Apple, the consumer company, quietly rehearsing a return to the one market it abandoned.
Sources
- [1] https://www.theinformation.com/articles/apple-considers-return-server-market-talked-nvidia-use-network-tech
- [2] https://www.bloomberg.com/news/articles/2026-09-16/apple-is-developing-enterprise-server-for-ai-age-report-says
- [3] https://the-decoder.com/apple-is-reportedly-building-an-enterprise-ai-server-with-its-own-m8-ultra-chips/
- [4] https://www.tomshardware.com/tech-industry/artificial-intelligence/apple-eyes-nvidia-nvlink-to-power-its-new-custom-m8-ultra-ai-servers-historically-bitter-rivals-reportedly-team-up-for-2029-data-center-push
- [5] https://www.macrumors.com/2026/09/16/apple-may-return-to-server-market/