1,024 INT8 MACs Inside the Shader Cores: Arm's Mali G2-Ultra NX Is the First AI-Native Mobile GPU
Arm's CSS for Mobile 2 embeds neural accelerators directly inside the Mali G2-Ultra NX shader cores — 1,024 INT8 MACs per clock, a seven-generation ISA overhaul, and DLSS-style neural upscaling promising 4x efficiency, targeting 2027 flagships like Samsung's Galaxy S27 and Xiaomi's XRING O3.
While the AI world spent the morning arguing about whether GPT-6 Astra really marks the arrival of AGI, Arm shipped a quieter but arguably more concrete piece of the AI era: on September 8, the company introduced Arm CSS for Mobile 2, its new AI-native compute platform for premium smartphones, headlined by the Mali G2-Ultra NX — the first Mali GPU with neural accelerators built directly into the shader cores.
The significance is architectural, not incremental. For two decades the smartphone has been a federation of separate specialists: a CPU for logic, a GPU for pixels, and — since around 2017 — a discrete NPU for neural networks, each with its own memory pathways and control overhead. Arm’s bet is that the next generation of mobile AI, from DLSS-style game upscaling to always-on agentic assistants, runs best when the neural math happens inside the graphics pipeline itself, sharing the GPU’s memory system, coherent caches, and control structures. The result is less data movement — the traditional tax of every mobile SoC — and a shorter path from model to pixel.
What shipping today actually contains
The announcement bundles three coordinated pieces of intellectual property:
- Mali G2-Ultra NX GPU — the first AI-native Mali, with dedicated matrix accelerators embedded in the shader cores, a new execution engine, and third-generation hardware ray tracing.
- Arm C2-Ultra and C2-Pro CPUs — new Armv9.3-A cores supporting SME2, the scalable matrix extension that brings matrix math to the CPU itself; Arm claims up to 1.7x higher AI-model performance versus the previous generation.
- The CSS for Mobile 2 platform — the compute subsystem that stitches compute, memory, software, and implementation optimizations into a licensable whole for silicon partners.
The numbers Arm is quoting are aggressive. Neural rendering on G2-Ultra NX can deliver up to 4x higher FPS and 4x higher performance efficiency versus native rendering, with up to 70 percent lower external memory traffic, according to Arm’s internal comparisons using the Neural Dawn demo developed with Sumo Digital. The new execution engine — the biggest Mali ISA overhaul in seven generations — contributes up to 24 percent higher benchmark performance and 14 percent higher non-AI gaming performance on its own, and third-generation ray tracing cuts DRAM traffic by up to 13 percent on leading RT benchmarks.
Independent analysis from Chips and Cheese adds the granular detail that Arm’s own materials omit. Each shader core’s matrix accelerator can do up to 1,024 INT8 MACs per clock or 512 INT16 MACs per clock, and it can clock up to twice the speed of the execution engines around it. The register file has grown by 25 percent, with per-warp register access doubled from 64 to 128 and allocation granularity improved to 16-register steps. Not every core needs the neural unit — the design mandates a minimum of six “NX cores” for Ultra branding — and silicon partners can choose how many to include. Xiaomi, whose XRING O3 already ships a G2-Ultra-class GPU (launched August 24), opted to give only half its shader cores the matrix accelerator; Chips and Cheese estimates the neural unit adds about 21 percent to a shader core’s area (~1.88 mm² vs ~1.55 mm²).
That flexibility cuts both ways. As Chips and Cheese notes, the matrix accelerator supports only INT8 and INT16 — no FP8, no BF16. For game upscaling and frame generation that quantization is fine; for running general LLM inference on-device, the omission is a genuine constraint.
Arm’s answer to DLSS — and a hedge on NPUs
The clearest way to understand what Arm is doing is to look at the three neural technologies that ride on the new hardware:
- Neural Super Sampling (NSS) — reconstructs higher-resolution images from lower-resolution renders, directly analogous to Nvidia’s DLSS.
- Neural Frame Rate Upscaling (NFRU) — generates intermediate frames from motion, depth, and rendered-frame data, the mobile counterpart to frame generation.
- Neural Super Sampling and Denoising (NSSD) — combines upscaling with denoising for demanding ray-traced scenes, the role DLSS Ray Reconstruction plays on desktop.
Videocardz frames it exactly that way: Arm is proposing “DLSS 5-like neural graphics for mobile GPUs with frame generation and ray reconstruction.” The difference is the business model. Nvidia owns both the silicon and the software stack for DLSS; Arm licenses IP to Qualcomm, MediaTek, Samsung, and Xiaomi, who each build their own chips. For neural graphics to reach “the billions of devices” Arm cites — more than 14 billion Mali GPUs shipped to date — the technique can’t live in a proprietary driver. It has to ship as portable IP.
That ecosystem effort started two years before the hardware. Arm’s Neural Technology program and an open development kit gave Tencent Games (Magic Dawn engine, an NSSD demo with Arena Breakout Infinite), Unity China (Tuanjie Engine), and the team behind Where Winds Meet early access, so engine-level integration landed before silicon. The Arm Neural Graphics Development Kit now provides Vulkan ML extensions, Unreal Engine plug-ins, an SDK for custom engines, and profiling and model-optimization tools.
The strategic read is subtler than “Arm built an NPU into the GPU.” Phone-makers already ship NPUs that excel at CNN-era vision workloads and increasingly at transformer inference. What NPUs don’t do is graphics. By putting matrix math where the pixels are, Arm positions the GPU as the home for neural graphics specifically — while the companion C2-Ultra CPUs with SME2 handle agentic-AI workloads that benefit from CPU-side matrix math. If the fourth wave of AI compute demand (chatbots → reasoning → agentic coding → computer use, as analyst Tae Kim framed it this week) really does land on phones, Arm wants every engine in the SoC matrix-capable, not just the NPU.
The fine print on the CPU claims
The C2-Ultra CPU story deserves its own skepticism. Arm claims a 15 percent peak performance uplift over C1-Ultra (12 percent average), but Chips and Cheese’s analysis of the endnotes shows the comparison rests on an 8.5 percent higher clock and a larger L2 cache on the FPGA-simulated C2-Ultra platform; at equivalent clocks the average uplift shrinks to roughly 3.2 percent, with a 7 percent peak IPC gain confirmed by Arm after the article’s publication. The claimed 38 percent power reduction, meanwhile, factors in process-node and implementation gains alongside microarchitecture. None of this makes C2-Ultra a bad core — it makes it an honest iteration in a generation whose real innovation budget clearly went to the GPU.
What lands in your hands, and when
CSS for Mobile 2 is IP shipping to partners now, which means phones in roughly a year. Notebookcheck’s coverage already places the architecture in Samsung’s Galaxy S27 and Vivo’s X500 Pro roadmaps, where it will compete with Qualcomm’s Oryon-core Snapdragon 8 Elite Gen 6. Xiaomi’s XRING O3 offers a preview today, shipping a G2-Ultra-class GPU with matrix accelerators in half its shader cores.
The pragmatic takeaway for developers: if you build mobile games or graphics engines, the Arm Neural Graphics Development Kit is available now, and the studios that integrated early — Tencent, Unity China, the Where Winds Meet team — will ship the first production titles. For everyone else, note the pattern: neural computation is dissolving into every compute engine Arm licenses. The question is no longer whether your phone has an AI accelerator. It’s how many, and which one runs the workload you care about.
Sources
Sources
- [1] https://newsroom.arm.com/blog/arm-mali-g2-ultra-nx-ai-native-mobile-graphics
- [2] https://chipsandcheese.com/p/arms-c2-ultra-g2-ultra-nx-and-css
- [3] https://www.androidauthority.com/arm-mali-g2-ultra-deep-dive-3706351/
- [4] https://videocardz.com/newz/arm-proposes-dlss-5-like-neural-graphics-for-mobile-gpus-with-frame-generation-and-ray-reconstruction