Intel Details Xeon 7 'Diamond Rapids': 256 Cores, 1.28 GB of Cache, and a 2027 Date
At Hot Chips 2026, Intel revealed its most modular Xeon ever — up to 256 P-cores across 22 chiplets with FP8 matrix acceleration, 16 DDR5 channels, and a fan-out fabric that finally answers AMD's chiplet playbook.
At Hot Chips 2026 this week, Intel pulled back the curtain on Xeon 7, codenamed Diamond Rapids — the most sophisticated server CPU the company has ever built, and the part it hopes will finally answer seven years of AMD chiplet dominance in the data center. The headline: the flagship SKU now reaches 256 performance cores in a single socket, up from the 192 cores originally planned, backed by a massive 1.28 GB of last-level cache, 16 channels of DDR5, and PCIe 6.0 throughout.
There is one catch, and it is a familiar one for Intel watchers: the chip will not ship until 2027.
From 192 cores to 256 — and why
Diamond Rapids was originally slated to launch in 2026 with up to 192 cores. That roadmap slipped, and Intel used the extra time to push the flagship configuration higher. According to ServeTheHome’s live coverage of the Hot Chips session, the 256-core configuration appears to have been added to the roadmap only recently — a direct response to AMD’s sixth-generation EPYC “Venice,” which will also top out at 256 cores (albeit with SMT enabled, for 512 threads).
The delay also reshaped the product line in another way. Intel canceled the eight-channel Diamond Rapids-SP variants earlier this year, leaving only the high-end AP-class parts. That makes Xeon 7 an HPC-centric, high-core-count family rather than a mainstream server platform — a narrower fight, but one Intel believes it can win on specs.
22 chiplets, three roles
The package is where Diamond Rapids gets genuinely interesting. It comprises up to 22 chiplets organized into three groups:
- 16 core compute dies built on Intel’s bleeding-edge 18A-P process — the refined 2nm-class node that debuted with Panther Lake. Intel claims 18A-P delivers the same performance at 18% less power, or up to 9% higher performance at iso-power, before counting the new cores’ microarchitectural gains.
- 4 compute building block (CBB) base dies on Intel 3-T, each carrying 320 MB of last-level cache. Up to four 16-core chiplets are stacked on each base die using Foveros 3D direct hybrid bonding, with a 3D crossbar linking cores to the shared LLC.
- 2 scalable fabric hub (SFH) dies on Intel 3 that centralize memory controllers and I/O — conceptually similar to the I/O dies AMD has shipped since EPYC Rome in 2019.
The Register’s deep dive calls this Intel’s “Rome moment”: after three generations of awkward tile experiments (Sapphire Rapids, Emerald Rapids, Granite Rapids), Intel has finally disaggregated compute from I/O and memory cleanly, with chiplets that can be added, subtracted, and reused per SKU.
Fan-out fabric and the UMA decision
The most technically consequential choice is what connects it all. Intel skipped its own EMIB 2.5D silicon-bridge packaging in favor of a UCIe-S-style interconnect carried over the organic package substrate — what Intel brands a “fan-out-fabric.” The reason is memory topology: EMIB’s edge-to-edge connectivity would have forced extra hops between the compute blocks and the fabric hubs. With fan-out fabric, each CBB talks directly to both SFH dies, letting Xeon 7 present as a uniform memory access (UMA) system by default — unlike Granite Rapids-AP, which exposed three NUMA nodes out of the box and forced software to do the balancing.
Coherence moved on-die as well. Intel replaced the DRAM-resident directory with an on-die snoop filter, preserving full ECC protection while cutting latency and coherence traffic. Each fabric hub contributes 16 MB of I/O cache, and CXL memory is supported in one-level or flat two-level modes with mirroring and per-path encryption.
The spec sheet
The resulting flagship numbers:
- Up to 256 P-cores (no hyperthreading — SMT is gone entirely, returning with the successor Coral Rapids)
- 1.28 GB of last-level cache (4 × 320 MB CBBs)
- 16 DDR5 channels at 8,000 MT/s, or 12,800 MT/s with Gen 2 MRDIMMs — up to 1.6 TB/s of memory bandwidth
- 128 lanes of PCIe 6.0 / CXL 3.0 / UPI 3 (four ×16 ports per fabric hub), plus 8 lanes of PCIe 4.0 for ancillary I/O
- TDP expected around 600 W, in line with AMD’s Venice flagship
On the ISA front, Xeon 7 is one of the first parts with full AVX 10.2 support, and the updated AMX matrix extensions now add FP8 — meaning low-precision AI inference can run directly on the CPU without a discrete accelerator. Intel is also adding “spill-and-fill” optimization instructions that expand the integer register file from 16 to 32 GPRs with new three-operand encodings, fully compatible with existing x86 software via recompilation.
Why no hyperthreading matters (and when it doesn’t)
Dropping SMT is the chip’s most controversial omission. In thread-happy throughput workloads, AMD’s 256-core/512-thread Venice could hold a double-digit percentage advantage regardless of whose cores are faster. But Diamond Rapids is explicitly an HPC part, and many HPC codes — including the HPL Linpack benchmark used to rank the Top500 — see little benefit from SMT, or actively regress. The Register also speculates Intel may position the part for low-latency agentic sandboxes: containers where AI-generated Python is executed and compiled, a workload where predictable single-thread latency beats logical core count.
The competitive picture
Spec-for-spec, Diamond Rapids lines up surprisingly well against Venice: both offer 256 cores, 16 memory channels at matching speeds, a gigabyte-plus of L3, similar I/O, and ~600 W TDPs. AMD’s structural advantages remain SMT and shipping first — Venice lands in 2027 as well, but AMD’s momentum in server share is unbroken, and Intel will spend the next year selling Granite Rapids and Clearwater Forest against AMD’s current generation.
Still, the architecture itself marks a genuine inflection. Moving coherence on-die, stacking 3D, and stepping to 18A-P close the efficiency gap that has plagued Intel since 2019. If Intel can price Xeon 7 aggressively — the way AMD’s Rome once undercut Intel with more cores per dollar — Diamond Rapids could make 2027 the first genuinely competitive server CPU fight in years. And in a data center market increasingly defined by AI buildouts, where every watt and every memory channel counts, that competition is arriving not a moment too soon.
Sources
- [1] https://www.theregister.com/hpc/2026/08/25/intel-diamond-rapids-xeon-7-cpu-deep-dive/5292427
- [2] https://www.servethehome.com/intel-diamond-rapids-the-2027-intel-xeon-at-hot-chips-2026/
- [3] https://www.tomshardware.com/pc-components/cpus/intel-xeon-7-diamond-rapids-comes-with-up-to-256-p-cores-1-28-gb-of-last-level-cache-next-gen-18a-p-cpu-also-brings-avx-10-2-and-uses-ucie-s-instead-of-emib
- [4] https://wccftech.com/intel-diamond-rapids-xeon-7-cpus-256-p-cores-1-28-gb-of-cache-12800-mtps-memory/
- [5] https://newsroom.intel.com/client-computing/intel-outlines-architectures-for-agentic-ai-at-hot-chips-2026