Apple's M6 and M5 Ultra: First 2nm Chip and Quad-Die Monster Aim squarely at On-Device AI
Apple's M6 is its first 2nm chip with a Dual 16-core Neural Engine, while M5 Ultra fuses four dies into the most powerful M-series SoC ever — 512GB unified memory, 1.2TB/s bandwidth, built to run frontier AI models locally.
Cupertino has picked its side in the AI hardware war. On August 25, 2026, Apple introduced two new system-on-chips that frame the same argument from opposite ends: that the next generation of AI should run on your desk, not someone else’s data center. The M6 — Apple’s first chip built on a 2-nanometer process — debuted inside a refreshed Mac mini, while the M5 Ultra, the most powerful M-series chip the company has ever shipped, landed in the new Mac Studio. Together they mark Apple’s most aggressive silicon push yet into on-device AI compute.
The timing is not accidental. The last-generation Mac minis became scarce earlier this year precisely because demand for machines that could run AI agents locally outstripped supply — an unusual problem for a desktop that had previously been pitched at everyday users. With M6 and M5 Ultra, Apple is leaning into that demand rather than treating it as a surprise.
M6: the 2nm mainstream chip with two Neural Engines
The M6 is the volume play. Built on a state-of-the-art 2nm process, it packs greater transistor density into a smaller die, and Apple has spent the density dividend almost everywhere:
- A 12-core CPU — two more cores than M5 — arranged as 2 super cores, 4 performance cores, and 6 efficiency cores. Apple claims the world’s fastest single-threaded CPU core, up to 1.2x faster multithreaded performance than M5, and up to 2.4x faster than the M1.
- A 12-core GPU — also two cores up on M5 — with a Neural Accelerator in every core. Peak GPU compute for AI rises nearly 30 percent over M5 and more than 8x over M1, which Apple says translates directly into faster prompt processing when interacting with on-device LLMs.
- A Dual 16-core Neural Engine — the headline feature. Two discrete engines deliver up to 2x the peak compute of previous generations, and system frameworks can dispatch work to both simultaneously, so applications see faster model execution without rewriting anything.
- Up to 170GB/s of unified memory bandwidth, a 10 percent increase over M5 and 2.5x over M1, with support for up to 32GB of unified memory — enough to run useful on-device LLMs for what Apple calls “secure and private agentic tasks.”
In the Mac mini, Apple says M6 delivers up to 4x faster AI performance, 2x faster storage and graphics, and 40 percent faster CPU performance compared to the M4 generation it replaces. The machine ships with macOS 27 “Golden Gate” preloaded, with full compatibility for the new Siri AI and Apple Intelligence features.
M5 Ultra: four dies, one chip, 512GB of memory
If M6 is the argument for everyone, M5 Ultra is the argument for the people who train and tinker. It is Apple’s first quad-die architecture: next-generation UltraFusion connects two dual-die M5 Max chips into a single logical processor, pushing inter-die bandwidth past 4.4TB/s with over 6x the connection density of the previous generation. Apple says the four dies behave as one unified processor — and the specifications read like a deliberate answer to workstation-class AI hardware:
- Up to a 36-core CPU (12 super cores, 24 performance cores): 1.25x higher single-threaded and 1.3x higher multithreaded performance than M3 Ultra.
- Up to an 80-core GPU with a Neural Accelerator per core — up to 4.5x the peak GPU compute for AI versus M3 Ultra and over 6x versus M1 Ultra, plus second-generation Dynamic Caching, hardware mesh shading, and third-generation ray tracing for up to 40 percent faster graphics.
- A 32-core Neural Engine driving Apple Intelligence and complex AI tasks on device.
- Up to 512GB of unified memory feeding a staggering 1.2TB/s of bandwidth — 50 percent more than M3 Ultra.
That memory subsystem is the whole point. Apple explicitly positions M5 Ultra as a machine for “running compute-intensive frontier AI models on device,” letting users keep huge datasets entirely in local memory, raise tokens-per-second throughput, and “run huge LLMs with hundreds of billions of parameters entirely on device.” A 512GB unified memory pool is oversized for almost any traditional desktop workload — but it is exactly the footprint you need to hold a deepseek-scale or larger model in memory with a generous KV cache. Apple even names the workflow: researchers can run local models in LM Studio Bionic to trigger complex simulations in MATLAB, and creators can generate AI images locally in apps like Draw Things.
Pricing and availability
The new Mac mini starts at $899 with M6 (16GB / 256GB) and $1,699 with M5 Pro (24GB / 512GB, configurable up to 18 CPU / 20 GPU cores, 307GB/s bandwidth) — both $100 more than the M4 generation’s starting prices. The Mac Studio starts at $2,499 with M5 Max (18-core CPU, up-to-40-core GPU, up to 128GB memory) and $5,499 with M5 Ultra. Pre-orders opened today; machines ship September 22, except the 512GB M5 Ultra Studio, which arrives in late October.
Both minis also get practical upgrades: storage that is twice as fast, 2.5-gigabit ethernet standard (10-gigabit optional), and Apple’s N1 wireless chip bringing Wi-Fi 7 and Bluetooth 6.
The strategy: own the edge of AI
Read together, the two launches sketch a coherent strategy. M6 brings meaningful AI compute to the price point where most people buy — two Neural Engines and GPU neural accelerators in an $899 box. M5 Ultra removes the last conventional excuse for sending frontier-scale models to the cloud: memory capacity and bandwidth. And the software stack — Core AI, Core ML, Metal, Xcode, Apple Foundation Models, and App Intents — is built to route work across CPU, GPU, and Neural Engine automatically, letting developers run and fine-tune large models locally.
There are caveats worth stating plainly. Apple’s comparative claims are against its own prior generations, not against current discrete-GPU workstations or dedicated AI accelerators, and real-world throughput for large-model inference depends as much on software maturity as on bandwidth. Raw training of frontier models will remain a data-center activity. But for inference, fine-tuning, and agentic workloads where privacy or latency matters, the pitch is now concrete: a 512GB, 1.2TB/s desktop that keeps the model — and the data — local.
Apple has spent six years making the case that unified memory architecture was an AI strategy in disguise. With M6 and M5 Ultra, the disguise is off.