← All posts / Tools

128 Cores on N3P and a Fifth of a Watt Saved per MIPS: Arm's Neoverse CSS N4 'Ranger' Rewrites the Semi-Custom Server Playbook

Arm's Neoverse CSS N4 'Ranger' compute subsystem scales from 8 to 128 N4 cores per die on TSMC N3P with LPDDR6, 256MB of L3 and 128 lanes of PCIe 7 — doubling socket performance over CSS N3 just as Arm announced Oracle and ByteDance's Volcano Engine are joining the AGI CPU ecosystem for agentic AI.

128 Cores on N3P and a Fifth of a Watt Saved per MIPS: Arm's Neoverse CSS N4 'Ranger' Rewrites the Semi-Custom Server Playbook

While most of this week’s AI headlines were consumed by model launches, Arm quietly shipped one of the more consequential infrastructure announcements of the month. On September 8, alongside its “Arm Everywhere” event in China, the company introduced the Neoverse CSS N4, codenamed Ranger — its most configurable Compute Subsystem yet, packing up to 128 cores per die on TSMC’s N3P process. In the same announcement, Arm confirmed that Oracle Cloud Infrastructure and ByteDance’s Volcano Engine are joining OpenAI, Meta, Cloudflare, SAP and Lenovo in building around its AGI CPU.

For anyone tracking where the compute for the agentic AI era actually comes from, this is the story under the story.

What Arm actually announced

A Compute Subsystem (CSS) is not a chip you can buy. It is Arm’s semi-custom program: a validated, integrated platform of CPU cores, interconnect, cache hierarchy, memory controllers and I/O that partners configure — core counts, cache sizes, connectivity — before taping out their own silicon on top of it. It is the same machinery behind the CPUs running in Azure, AWS Graviton and Google Axion fleets today.

Neoverse CSS N4 is the N-series (efficiency-optimized) successor to CSS N2, and the spec sheet reads like a direct response to what agentic workloads are doing to data-center CPUs:

  • 8 to 128 Neoverse N4 cores per die (the cores are codenamed Dionysus), clocking up to 3.8 GHz — with the caveat that peak clocks presumably drop as core counts rise
  • Up to 256 MB of shared L3 cache per die, up from 64 MB on CSS N2
  • Up to 2 MB of L2 per core (double CSS N2’s 1 MB), plus 64 KB L1 instruction and 64 KB L1 data caches per core
  • DDR5 or LPDDR6 memory support — LPDDR6 being a first for the platform, and a meaningful option for throughput-per-watt-bound deployments
  • Up to 128 lanes of PCIe 7/6 and CXL 4.0, up from 64 lanes of PCIe 5
  • Multi-chiplet and multi-socket scaling beyond the 128-core-per-die ceiling, with UCIe die-to-die interconnects and what Arm calls “partner-specific PHYs”

Against CSS N3, Arm’s claimed numbers are 2x socket-level performance, 1.25x performance per watt, and 1.75x memory bandwidth. Tom’s Hardware notes those figures were quoted for a 128-core part running at 3 GHz with the full 2 MB of L2 per core.

Why the N-series matters now

Arm’s Neoverse lineup splits between the V-series (maximum performance — the lineage inside Nvidia’s Grace and AWS Graviton) and the N-series (performance per watt). Historically, N-series cores landed in infrastructure processors and efficiency-tier cloud instances — think Azure Cobalt 100 (N2) or Intel’s IPU Adapter E2100 (N1), not flagship server CPUs.

That positioning is exactly why CSS N4 is interesting. Arm’s own framing in the announcement is that agentic AI — agents that reason, retrieve, call tools, query databases and talk to other agents — is shifting meaningful computation away from the accelerator and onto general-purpose CPUs. Scale-out data-plane serving wants throughput efficiency; that is N-series territory. Arm claims Arm-based rack-scale servers have already overtaken x86 as the dominant accelerated computing platform, and it is now offering two on-ramps: build your own silicon on CSS N4, or buy production-ready AGI CPU silicon built on V3 cores.

The AGI CPU side of the announcement carries its own weight. Arm disclosed that Volcano Engine — ByteDance’s cloud division — is enabling next-generation cloud services on the platform and bringing “the first agentic sandboxes powered by Arm AGI CPU” to market, and that Oracle Cloud Infrastructure has joined the ecosystem. Add the previously named OpenAI, Meta, Cloudflare, SAP and Lenovo deployments, and what started as an IP licensing story has become a silicon platform story with a genuinely short customer list of the largest compute buyers on earth.

The fine print

A few honest caveats belong next to the headline numbers. First, CSS N4 is an announcement for silicon builders, not buyers — Arm has not named any CSS N4 customers yet, and N-series parts typically reach the wild a design cycle after the platform announcement. Traditionally only a few large CSS contracts are needed for the platform to matter, but those contracts haven’t been disclosed.

Second, the 2x/1.25x/1.75x figures are vendor claims against Arm’s own previous generation, quoted at a specific configuration. Third, the efficiency tier still has to prove itself against the x86 status quo in agentic serving stacks — a workload profile where memory bandwidth and I/O lanes (both doubled here) may matter more than raw core counts.

And there is competitive context: Arm’s 2024 roadmap already promised a CSS V4 (codenamed Vega) alongside this N4 generation, so Ranger is half of a two-pronged platform refresh rather than the whole story.

Why it matters

The AI infrastructure narrative has spent two years obsessing over GPUs. The quieter trend is that once models stop being chat interfaces and start taking actions, the CPU fleet underneath the accelerators has to grow: more orchestration, more data movement, more concurrent processes per rack. Every major hyperscaler is now either building custom Arm silicon or buying Arm-based production chips, and Arm’s answer to that demand is to make both paths cheaper — custom silicon with the fastest CSS-to-silicon path it has ever offered, and off-the-shelf AGI CPUs for everyone else.

For the semiconductor industry, CSS N4 is also a statement about process leadership: 128 efficiency cores on TSMC N3P with LPDDR6 and PCIe 7 is a 2027-ish product when it lands in partner silicon, and it will arrive into a market where “performance per watt per agent” is starting to appear in procurement conversations.

The agentic era, it turns out, is partly a CPU sales story. Arm just made that explicit.