← All posts / Industry

Nvidia Tells Its Biggest Customers: AI Servers Will Cost 15%+ More Next Year

Bloomberg reports Nvidia's largest customers have been warned that AI servers — including Vera Rubin and Grace Blackwell systems — will cost more than 15% extra when they ship in early 2027, as soaring memory prices rewrite the economics of the AI buildout.

Nvidia Tells Its Biggest Customers: AI Servers Will Cost 15%+ More Next Year

The next bill for the AI boom just arrived, and it is bigger than almost anyone budgeted for. According to a Bloomberg News report published Saturday, some of Nvidia’s largest customers have been formally notified that prices for servers containing its AI chips will rise by more than 15% in many cases — and the increases will take effect on systems shipped early next year, including flagship machines built around the Vera Rubin and Grace Blackwell platforms.

Reuters, which syndicated the report, added the key mechanics: server makers that build systems under contract for large data-center operators — including Microsoft, Alphabet’s Google, and Oracle — have already passed the upcoming increases along to their own customers. The size of each hike will depend on the chip generation and memory configuration of the system in question. Nvidia declined to comment, and Reuters said it could not independently verify the figures, but the direction of travel is unmistakable: the era of perpetually falling per-unit AI infrastructure costs has, at least for now, come to an end.

The culprit is memory, not silicon

The single biggest driver, according to the report, is the exploding cost of memory chips. This is not a new story — it is a squeeze that has been building for over a year and has now reached the point where it moves the sticker price of entire racks.

The numbers behind the crunch are staggering. Server-grade DDR5 prices have been on a run that analysts expect to double year-over-year by late 2026. A 64GB DDR5 module that cost roughly $150 not long ago has been quoted around $500. Research firm Counterpoint forecast a 50% rise in overall memory prices through Q2 2026, with server-specific DRAM facing the steepest hikes. On the high-bandwidth side, CNBC reported in January that AI memory — particularly HBM stacked around GPUs like those in Nvidia’s NVL72 racks — was effectively sold out, with supply spoken for well in advance.

The physics of the problem is simple. AI data centers now absorb an estimated 70% of global memory chip output, and every new GPU generation is hungrier than the last: Vera Rubin-class systems pair each compute die with more HBM and more LPDDR5X than the generation before. SK Hynix has warned the shortage may persist beyond 2030. When three suppliers — Samsung, SK Hynix, and Micron — control essentially all advanced DRAM and HBM capacity, and hyperscalers bid against smartphone and PC makers for the rest, prices have only one way to go.

What it means for the hyperscalers

For Microsoft, Google, Oracle, Meta, and the swelling ranks of “neocloud” providers, a 15%+ increase on AI servers is not a rounding error — it is a direct hit to the capital expenditure lines that define the entire AI investment cycle. Nvidia, whose chips underpin the majority of the AI infrastructure buildout, has become the proxy for the whole ecosystem, and its fiscal calendar now functions as a quarterly referendum on AI demand.

The timing is pointed: Nvidia reports second-quarter results on August 26, and this pricing news lands days before. Investors will now be watching not just data-center revenue but gross-margin commentary for signs of how much of the memory inflation Nvidia can pass through versus absorb. Historically, Nvidia has had the pricing power to push component inflation downstream — that is precisely what this notification appears to be. The open question is whether cloud providers, already committing hundreds of billions of dollars to multi-year buildouts, will push back, re-time purchases, or simply pay.

There are second-order effects worth tracking. If server BOM costs keep climbing, cloud GPU rental rates face upward pressure just as providers have been cutting per-token prices to win AI workloads. Morgan Stanley and other analysts have warned the peak of cloud rate hikes may arrive in the second half of 2026. Enterprises signing multi-year AI capacity contracts this quarter may want to negotiate price-escalation clauses now rather than discover them in 2027.

The ripples beyond the data center

The memory crisis is not contained to AI racks. Consumer graphics cards have already seen repeated increases — one widely circulated report tallied a 20–30% hike across Nvidia’s consumer lineup as its third major increase of 2026. Smartphone and laptop makers are absorbing double-digit component cost increases that SK Hynix and others have flagged, and some of that will inevitably reach retail prices. The Bloomberg piece quotes an industry participant acknowledging the obvious: “Like others, we are seeing increased input costs driven primarily by the rising prices of DRAM and NAND.”

Meanwhile the memory makers themselves have become the quiet winners of the cycle. Samsung, SK Hynix, and Micron each crossed a trillion dollars in market value within weeks of each other earlier this year — an outcome few would have predicted when DRAM was a commodity business with razor-thin margins. Samsung reclaimed the No. 1 spot in global DRAM revenue in Q2 2026 with a 39% share, while Micron narrowed the gap and China’s CXMT posted its strongest showing yet.

The takeaway

For a year, the AI infrastructure story has been about compute scarcity. The Bloomberg report makes official what the supply chain has signaled for months: memory is now the binding constraint, and it is expensive enough to move the price of an entire server by double digits. Nvidia is passing the cost through while its pricing power lasts; the hyperscalers will pass it to the cloud; and eventually it lands on anyone who buys a token, a GPU-hour, or a laptop.

Whether this is a painful blip or the beginning of a structurally more expensive era of computing depends on memory supply elasticity — new fab capacity, HBM4 yields, and whether players like CXMT can scale into advanced nodes fast enough to matter. SK Hynix’s warning that the shortage could last beyond 2030 suggests the industry is not betting on quick relief. Either way, the assumption that AI capability gets monotonically cheaper per dollar of hardware just took a 15%+ haircut.

Nvidia is expected to address pricing and demand when it reports fiscal Q2 results on August 26.