128 Cores, 3.8 GHz, LPDDR6: Arm's Neoverse CSS N4 Is the Semi-Custom Blueprint for the Agentic AI Data Center
Arm's Neoverse CSS N4 doubles the per-die core ceiling to 128, adds LPDDR6 and PCIe Gen 7, and claims 2x socket performance and 1.25x performance per watt over CSS N3 — a semi-custom on-ramp aimed squarely at agentic AI's CPU-hungry workloads.
On September 8, 2026, Arm pulled back the curtain on Neoverse CSS N4 — codenamed Ranger — its most configurable Compute Subsystem to date and the company’s clearest statement yet about where it thinks AI infrastructure is heading: away from a single accelerator-centric design point and toward a swarm of semi-custom host CPUs orchestrating agents at scale.
The headline numbers are straightforward. CSS N4 scales from 8 to 128 N4 cores per die, fabricated on TSMC’s N3P process, with clock speeds reaching 3.8 GHz — a first for Arm’s N-Series subsystems. Each core carries 2 MB of L2 cache, with up to 256 MB of shared L3 on the die. Against the previous-generation CSS N3, Arm claims up to 2x socket-level performance, 1.25x performance per watt, and 1.75x memory bandwidth.
But the raw specs only tell half the story. The strategic shift is in what Arm chose to connect: LPDDR6 memory support (another first for a Neoverse CSS) and PCIe Gen 7 with up to 128 lanes, alongside CXL 4.0 and coherency options. As TechTimes noted, 128 lanes of PCIe Gen 7 is enough to hang eight x16 accelerator links off a single N4-based chip — turning the host CPU into a genuine orchestration hub for dense AI clusters rather than a conveyor belt feeding a GPU.
Why the agentic era is a CPU story
The most interesting part of Arm’s announcement isn’t silicon at all — it’s the workload argument. As AI moves “from inference to action,” agents reason, retrieve data, call tools, and interact with databases and other agents. Every one of those steps generates CPU-bound work: JSON parsing, sandbox execution, key-value lookups, network round-trips, scheduler overhead. The GPU still does the heavy lifting of token generation, but the surrounding choreography lands squarely on the host CPU.
Arm’s positioning reflects an industry-wide realization that the host-CPU layer, long treated as a commodity in AI servers, is becoming a first-order bottleneck — and a differentiation opportunity. Arm now offers two paths on a common platform: CSS N4 for customers building their own differentiated silicon (hyperscalers, DPU vendors, networking specialists), and the Arm AGI CPU for those who want production-ready chips tuned for responsive agentic workloads.
The ecosystem claims around AGI CPU read like a who’s-who list: OpenAI, Meta, Cloudflare, Oracle, SAP, Lenovo, Supermicro, and Verda are all described as developing solutions around it. Google Cloud is running agent sandboxes on Axion-based GKE; Microsoft Azure is accelerating sandbox tool execution with Cobalt 200; NVIDIA’s Vera is designed for agentic workloads; and ByteDance’s Volcano Engine is bringing the first AGI-CPU-powered agentic sandboxes to market.
The semi-custom pitch
Compute Subsystems have always been Arm’s answer to a hard problem: custom silicon is expensive and slow, but off-the-shelf designs leave performance on the table. A CSS is a pre-integrated, pre-validated platform — cores, interconnect, memory controllers, I/O — that partners can configure around their own requirements. Arm says CSS N4 is its “most configurable CSS yet” with the “fastest path from CSS to silicon,” supporting both single-die and multi-chiplet designs targeted at cloud servers, networking equipment, DPUs, and specialized accelerators.
The configurability range is unusually wide: an 8-core die for a lean DPU and a 128-core throughput monster can share the same software foundation, IP validation, and Total Design ecosystem partnerships. Arm also says it is extending the Total Design collaborative model — which pulls in EDA, foundry, and design-services partners for earlier IP validation — into physical AI, hinting at robotics and edge-adjacent designs to come.
Context: Arm’s data-center momentum is real
Arm backed its launch with a striking datapoint: according to IDC, Arm-based rack-scale servers have overtaken x86 as the dominant accelerated computing platform. That claim deserves scrutiny — “accelerated computing” is a specific segment, and x86 remains dominant across the installed base — but it reflects a genuine inflection. AWS Graviton, Google Axion, Microsoft Cobalt, and NVIDIA Grace/Vera have collectively moved Arm from exotic to mainstream in the span of a few years, and agentic workloads are exactly the kind of bursty, many-process, I/O-heavy load where Arm’s performance-per-watt advantage compounds at rack scale.
The competitive backdrop matters too. Intel and AMD are not standing still — Granite Rapids and Turin-class x86 parts push core counts upward, and both camps are chasing the same efficiency narrative. Arm’s bet is that in the agentic era, the winning metric shifts from peak throughput per socket to responsive performance per watt: how quickly can a host CPU wake an agent sandbox, service a tool call, and stream data to an accelerator without stalling it. CSS N4’s 3.8 GHz ceiling (high for a 128-core efficiency-leaning part) and LPDDR6 support are aimed directly at that latency-bandwidth trade-off.
What to watch
Three open questions follow the launch. First, adoption timing: CSS N4 is a design platform, not a shipping chip — customer silicon is typically 18–30 months out, meaning N4-based parts will land in the 2028 timeframe, competing against whatever x86 and RISC-V offer by then. Second, software gravity: the common Neoverse software foundation is Arm’s moat, but agentic frameworks evolve monthly, and the AGI CPU’s “production-ready” claim will be tested by real deployment diversity. Third, the power ceiling: LPDDR6 saves power on the memory side, but 128 cores at 3.8 GHz on N3P will demand serious packaging and cooling engineering in dense configurations.
For now, the signal is clear. The industry’s compute conversation is widening from “how big is the GPU” to “how well does the whole system orchestrate agents” — and Arm has just laid down its most aggressive marker yet on the CPU side of that equation.
Sources are listed in the article metadata.
Sources
- [1] https://newsroom.arm.com/news/arm-agi-cpu-neoverse-css-n4-agentic-ai
- [2] https://www.tomshardware.com/pc-components/cpus/arm-debuts-next-gen-semi-custom-neoverse-css-n4-ranger-platform-compute-subsystem-packs-up-to-128-cores-per-die-on-tsmc-n3p
- [3] https://www.arm.com/products/cloud-datacenter/neoverse-compute-subsystems/css-n4
- [4] https://www.techtimes.com/articles/327003/20260908/arm-neoverse-css-n4-doubles-core-ceiling-adds-lpddr6-pcie-gen-7-ai-servers.htm
- [5] https://aiweekly.co/alerts/arm-debuts-neoverse-css-n4-ranger-with-up-to-128-cores-at-38ghz-on-tsmc-n3p-for