A smaller entry to the DGX Spark line
NVIDIA has announced a 64GB version of DGX Spark, its compact local-AI system for developers and researchers. The company’s 2 October post says the configuration will be offered by Acer, ASUS, Dell, Gigabyte, HP and MSI from Friday 23 October, starting at US$4,999. It keeps the GB10 Grace Blackwell Superchip, DGX OS and the NVIDIA AI software stack used in the existing 128GB version. This is an announcement of an upcoming configuration, not a claim that every buyer can purchase one today.
The central proposition is to run models and agent workflows on premises rather than sending every experiment to a cloud instance. Local operation can be attractive when a team wants to work with private code or data, control the runtime directly or keep a development agent available continuously. It does not remove the need for security controls, software updates or sound handling of the data the agent reads.
Memory is the first design constraint
NVIDIA says a single 64GB system can support models with up to 100 billion parameters, depending on the model and workload. That headline does not mean every model of that size will run at a useful speed or with room for a long context window. Memory use depends on quantisation, activation and cache requirements, concurrent requests and the inference stack. Teams should test their own target model and agent pattern rather than buying solely from a parameter count.
The platform combines Grace Blackwell compute, unified memory and ConnectX-7 networking with CUDA-accelerated software. NVIDIA lists Agent Toolkit, CUDA-X libraries and Nemotron models alongside common runtimes such as Ollama, vLLM and PyTorch. That ready-to-use stack could shorten setup time for teams already working in NVIDIA’s ecosystem, while developers using a different framework should confirm support and performance for their own pipeline.
Two boxes, one larger working set
NVIDIA is also promoting Sync Cluster Assistant, which connects two DGX Spark systems with a QSFP cable through their built-in ConnectX-7 interfaces. It says the combination pools 128GB of memory, supports models up to 200 billion parameters and can reach up to 1.7 times the performance of one unit in a Qwen 3.8 27B test. The benchmark is a vendor-reported result for a particular workload, not a general guarantee that every application will scale by the same factor.
Cluster Assistant detects connected devices, checks their configuration and sets up the network. This addresses a real source of friction: multi-node inference can require more operational effort than a single local machine. Even with automated setup, however, distribution can add communication overhead. A useful pilot should compare one-node and two-node throughput, latency and total energy use on the actual models the team intends to run.
Agents and a launcher still to come
NVIDIA frames the new system around agentic work. It suggests keeping a coding or research agent running, serving a model to everyday laptops, and increasing capacity when a task outgrows one device. At the end of October, the company plans to add NVIDIA Sync Model Launcher, intended to download and launch Qwen3.8 27B on one system or a cluster and make it accessible from a laptop. That future tool should be evaluated once it ships; it is not a present capability of the 2 October announcement.
The post also points developers to playbooks for OpenClaw, NemoClaw, Hermes Agent and OpenShell, and mentions a forthcoming Blender installer. These show a broader local-AI ecosystem, but they serve different users. A developer running an autonomous code agent needs a different evaluation from a creator testing image tools or a researcher fine-tuning a model.
What to check before ordering
The 64GB price and 23 October date are starting points for planning. Buyers should confirm local currency, partner stock, warranty, power needs and the software version available on delivery. A two-unit deployment also requires two systems and the relevant networking, so its cost should be compared with a larger single machine or rented cloud capacity. The best option will depend on utilisation: hardware that runs daily has a different economics from an occasional experimental workload.
NVIDIA’s announcement makes local AI more accessible within its DGX line and simplifies a path to additional memory. The strongest case for it is a repeatable workflow where privacy, control or sustained use justifies local infrastructure. The claim to test is not merely whether a model launches, but whether the full agent can run reliably, safely and fast enough on the configuration the team can actually buy.
Developers should also think about the lifecycle after the first demo. A persistent local agent needs monitoring, permission boundaries and a way to recover from failed tasks or interrupted sessions. Model files and project data can consume substantial local storage, while an unattended tool-using agent may need restrictions on network access and shell commands. DGX Spark supplies compute and a software ecosystem, but those operating policies still belong to the team that deploys it.