Half the Memory, Two-Thirds the Price: DGX Spark 64GB Brings Local Agents Down to $4,999
NVIDIA's new 64GB DGX Spark cuts the entry price of a Grace Blackwell desktop to $4,999, runs 100B-parameter models on device, and clusters two units to 128GB via the new Sync Cluster Assistant.
NVIDIA has found a new way to say “local AI is getting cheaper”: cut the memory in half and knock a third off the price. On October 2, the company announced a 64GB configuration of the DGX Spark, its Grace Blackwell-powered desktop AI machine, which goes on sale Friday, October 23 through Acer, ASUS, Dell, Gigabyte, HP and MSI, starting at $4,999 — roughly $2,400 less than the 128GB model’s usual street price of around $7,499.
The pitch is straightforward. The 64GB SKU keeps everything that defines the platform — the GB10 Grace Blackwell Superchip, DGX OS, the ConnectX-7 networking, and the full NVIDIA AI software stack — and gives up only memory capacity. According to NVIDIA, that is enough to run models of up to 100 billion parameters entirely on device, along with the agentic applications built on top of them, privately and without any cloud dependency.
What the new SKU actually is
DGX Spark has always been an odd but compelling product: a compact, ~240W desktop that packages a Grace Blackwell compute chip, unified LPDDR5X memory, 200GbE ConnectX-7 networking, and a CUDA-accelerated software stack into something that sits next to your monitor instead of in a data center. It was positioned from the start as a “personal AI supercomputer” — a place to experiment with frontier-class open models and your own data without spinning up a cloud instance for every task.
The 64GB configuration is being sold exclusively through NVIDIA’s manufacturer partners rather than direct, which is presumably how the partners make the lower price point work. Tom’s Hardware characterized the move as a lifeline for local-AI enthusiasts “amid the rampocalypse” — a nod to the memory supply squeeze that has made high-capacity LPDDR5X configurations both expensive and scarce over the past year. For buyers whose workloads fit in 64GB, the new SKU is simply the same machine for materially less money.
The headline capability number: up to 100-billion-parameter models on a single unit. That comfortably covers today’s most popular open-weight workhorses, and NVIDIA says the system ships ready for agent development from day one — the NVIDIA Agent Toolkit, CUDA-X AI libraries, Nemotron open models, and runtimes like Ollama, vLLM, llama.cpp, LM Studio and PyTorch with CUDA are all supported out of the box. Blender is also among the first major creator applications to support the platform, with a prebuilt installer coming soon.
The cluster story: two Sparks, one brain
The more interesting part of the announcement is not the memory cut — it is the software that ships alongside it. NVIDIA Sync Cluster Assistant is designed to make scaling from one DGX Spark to two nearly frictionless. Connect two units directly with a QSFP cable and they pool their memory to 128GB, expand model support to up to 200 billion parameters, deliver twice the memory bandwidth, and in NVIDIA’s testing with Qwen 3.8 27B, delivered up to 1.7x the performance of a single system.
The Sync app handles the parts that normally make clustering painful: it detects connected units, validates device configuration, and sets up the ConnectX-7 network automatically. Every node runs the same NVIDIA software stack, so nothing needs to be reconfigured when going from one machine to two. In practice, this turns the 64GB SKU into a building block — start with one, add a second when your models outgrow 64GB, and treat the pair as a single 128GB machine.
By the end of the month, NVIDIA also plans to ship Sync Model Launcher, which the company describes as making local AI “as simple as clicking a few buttons.” It will download and launch models like Qwen3.8 27B on a single system or a cluster, configure the model to run across connected devices, expose it to the user’s laptop, and even set up OpenCode so developers can start coding against the model in a browser.
Why this matters
The economics of local inference have been improving on two fronts: open models keep getting more capable per parameter, and the hardware to run them keeps getting cheaper per gigabyte of usable memory. A $4,999 box that runs 100B-parameter models with no subscription and no data leaving the building is a meaningful data point in the running argument between local-first and cloud-first AI — particularly for developers running always-on coding or research agents, where per-token cloud costs compound quickly.
NVIDIA sketched three concrete workflows for the new configuration: running a coding or research agent around the clock (with a cluster adding capacity for larger models, longer contexts, or multiple concurrent agents); offloading model inference from an everyday PC to the Spark so the laptop stays free; and scaling a single task across two units when it outgrows one, without touching the software environment.
There is also a clear competitive backdrop. Apple’s M5 Ultra Mac Studio has been winning editor’s-choice comparisons against the 128GB Spark on memory-per-dollar grounds, and a wave of Windows “RTX Spark” PCs from Acer, ASUS, Dell, HP, Lenovo, Microsoft and MSI is arriving this month. The 64GB SKU, with its clustering path to 128GB and beyond, is NVIDIA’s answer to both: a lower entry price today, and an upgrade path that a sealed desktop from Cupertino doesn’t really offer.
The caveats
Half the memory is still half the memory. Anyone whose target models sit between 100B and 200B parameters will need either the 128GB unit or two clustered 64GB units — at which point the price advantage evaporates ($9,998 for the pair versus roughly $7,499 for a single 128GB machine, though with double the compute and memory bandwidth). LPDDR5X’s relatively modest bandwidth also means token generation speeds trail discrete-GPU setups for single-stream workloads. And the platform still lives on ARM64, which remains a friction point for some CUDA software despite a year of maturing support.
But as a starting point, the math is hard to argue with: the same Grace Blackwell platform, the same software stack, agent-ready out of the box, and a $2,400 lower bar to entry. For the local-AI community that has been priced out of the 128GB model — or simply tired of fighting for cloud GPU capacity — the 64GB DGX Spark lands October 23.
Sources
Sources
- [1] https://blogs.nvidia.com/blog/local-ai-dgx-spark-64gb-sync/
- [2] https://www.tomshardware.com/pc-components/gpus/nvidia-introduces-64gb-dgx-spark-to-throw-local-ai-fans-a-lifeline-amid-the-rampocalypse-new-gb10-config-starts-at-usd4999-for-those-who-can-work-with-less
- [3] https://www.nvidia.com/en-us/products/workstations/dgx-spark/
- [4] https://forums.developer.nvidia.com/t/nvidia-dgx-spark-64gb-gives-developers-more-ways-to-build-and-scale-local-ai/384838