NVIDIA is adding a 64GB unified-memory configuration of DGX Spark through hardware partners including Acer, ASUS, Dell, Gigabyte, HP and MSI. The system keeps the same GB10 Grace Blackwell Superchip and DGX software stack as the larger configuration.
NVIDIA says the 64GB version can run models with up to 100 billion parameters on the device, depending on the model and memory requirements. It can also be connected with other DGX Spark systems when a workload needs more memory and compute.
What DGX Spark is designed for
DGX Spark is aimed at developers, researchers and AI builders who want a local machine for model development and inference. The appeal is control: workloads can run on the device without sending every prompt or data file to a remote service.
Local AI does not mean every model will run well on every configuration. Model size, quantization, context length, memory use, and workload type all affect practical performance.
What the new 64GB configuration changes
The 64GB model gives developers another entry point into the DGX Spark platform. NVIDIA says it keeps the GB10 Grace Blackwell Superchip, DGX OS and the full NVIDIA AI software stack.
The main difference is memory capacity. Some larger models or long-context workloads need more memory than a 64GB system can comfortably provide. For smaller models and development workflows, however, 64GB may be enough.
Local AI can be useful for development
Developers often need to test prompts, model behavior, retrieval pipelines, agents, and data-processing workflows repeatedly. Sending every experiment to a hosted API can add cost and can create data-handling concerns.
A local system can provide a private development environment for supported models. It can also make experimentation easier when a team needs predictable access to a model without depending on an external API connection.
But local hardware has its own costs. Electricity, hardware purchase price, maintenance, storage, model downloads, updates, and developer time all count.
How to decide if 64GB is enough
Start with the exact models and workflows you want to run. Do not choose the hardware from a parameter-count headline alone.
- List the models you need.
- Check their memory requirements at the intended quantization.
- Include context and runtime overhead in the estimate.
- Measure tokens per second on your real workload.
- Check whether several services need to run at the same time.
- Leave headroom for the operating system and supporting tools.
A model that technically fits in memory may still be unpleasant to use if it leaves too little headroom or runs too slowly for the workflow.
When 128GB makes more sense
The larger configuration is more useful when a team needs bigger models, longer contexts, or several memory-heavy processes. It can also provide more room for experimentation before developers have to optimize or split a workload.
For a single developer learning local AI, the 64GB option may be a practical starting point if the target models fit. For a research team working with larger models, memory headroom can be worth paying for.
Clustering multiple DGX Spark systems
NVIDIA says multiple DGX Spark systems can be connected using its Sync Cluster Assistant for workloads that need more memory and compute.
That introduces a different design problem. Once a workload spans machines, networking and communication overhead matter. A model that fits on one system may behave differently when its work is distributed across several systems.
Teams should benchmark the full workflow rather than assuming that two machines automatically provide twice the useful performance.
Local AI versus cloud AI
Cloud APIs are often easier when the goal is quick access to many models, flexible capacity, and minimal hardware management. Local hardware becomes more attractive when privacy, offline access, repeated experimentation, or predictable local inference matters.
Many teams will use both. A local machine can handle development, testing and selected workloads, while cloud services can provide access to larger models or bursts of compute.
What developers should test first
If you are considering a local AI workstation, create a benchmark before buying. Use the same prompts and tasks you expect to run after the purchase.
Measure response speed, memory use, quality, power use, setup time, and total cost. Include the tools you actually use, not only a model benchmark.
For agent workflows, test tool calls and long-running tasks. For coding, test real repository tasks. For retrieval, load a representative document set and measure both answer quality and latency.
The practical value of the 64GB option
The new configuration makes local AI more flexible because developers can choose a memory tier that better matches their workload. The key is to treat memory as one part of the system rather than the only specification.
For many builders, the right question is not “Can this machine run a 100-billion-parameter model?” It is “Can this machine run the models and workflows I actually need at a useful speed and cost?”