NVIDIA DGX Spark 2025: A Compact AI Supercomputer Built for Local LLMs and AI Agents
The rapid rise of artificial intelligence has changed what people expect from a modern PC. Not long ago, the idea of running serious AI workloads locally sounded unrealistic for most creators, developers, researchers, and small teams. Today, the market is moving quickly toward compact systems designed specifically for AI development, local inference, model fine-tuning, agents, and data science.
The NVIDIA DGX Spark 2025 is one of the most interesting examples of that shift. It is a Mini PC, but calling it only a Mini PC does not really capture what it is designed to do. With a price of $4,699, DGX Spark is positioned as a personal AI supercomputer for users who want powerful local AI acceleration without relying entirely on cloud infrastructure.
At the heart of the system is NVIDIA’s GB10 Grace Blackwell Superchip, a compact but highly capable processor designed to bring technologies from large AI data centers into a desktop-friendly machine. The result is a small, power-efficient workstation that can handle large language models, AI agents, data science tasks, rendering, visualization, and compute-heavy development workflows.
NVIDIA DGX Spark and the rise of the local AI workstation
AI PCs are becoming a major focus across the technology industry, but not all AI PCs are built for the same purpose. Many consumer systems focus on lightweight AI features inside Windows applications, camera effects, productivity tools, and local assistant-style experiences. DGX Spark is different.
This system is aimed at people who want to build, test, run, and fine-tune AI models locally. It is not just about adding AI features to a normal desktop experience. It is about creating a compact development platform that can support serious AI workloads directly on your desk.
The DGX Spark was among the first systems to define this new category of personal AI boxes. Since then, more compact AI-focused machines have started appearing, but DGX Spark stands out because of how closely it connects to NVIDIA’s broader AI software and hardware ecosystem.
The machine is designed to let developers work locally, then move workloads to larger accelerated infrastructure when needed. That makes it especially appealing for AI engineers, researchers, startups, students, and enterprises that want a smaller system for prototyping before scaling to cloud or data center environments.
GB10 Grace Blackwell Superchip: The core of DGX Spark
The most important part of the NVIDIA DGX Spark is the GB10 Superchip. This chip brings the Grace Blackwell architecture into a compact form factor suitable for a Mini PC or workstation.
The GB10 is built using advanced multi-die packaging and is manufactured on TSMC’s 3nm process technology. It combines a CPU dielet and GPU dielet into one package, creating a unified design that can efficiently share memory and data across different parts of the system.
The chip includes two main sections. The S-dielet contains the CPU, memory subsystem, and related system components. The G-dielet contains the GPU core. Together, they allow DGX Spark to deliver high AI performance in a much smaller footprint than traditional workstation or server-class systems.
The goal is clear: bring key data center technologies into a compact desktop system. NVIDIA has included support for technologies such as NVFP4, CUDA, TensorRT, vLLM, NVLINK C2C, ConnectX-7 networking, unified memory, and other parts of its AI software and hardware stack.
CPU architecture and performance design
The CPU inside the GB10 Superchip is based on Arm architecture v9.2 and features 20 cores in total. These cores are split into two clusters of 10 cores each.
One cluster uses Cortex-X925 cores, designed for high performance. The other cluster uses Cortex-A725 cores, which are optimized for efficiency and balanced workloads. Each core has its own private L2 cache, while each cluster receives 16 MB of L3 cache, giving the CPU a total of 32 MB of L3 cache.
This CPU design gives DGX Spark enough general-purpose performance to handle development workflows, system tasks, data preparation, software environments, and AI pipeline management. However, the real star of the machine is the integrated Blackwell GPU.
Blackwell GPU performance for AI workloads
The GPU inside the GB10 Superchip is based on NVIDIA’s Blackwell architecture. Since it is integrated into the same package as the CPU, it can be considered an integrated GPU, but its capabilities are far beyond what most people associate with that term.
The Blackwell GPU includes 5th Generation Tensor Cores, RTX ray tracing cores, and support for DLSS 4. For AI workloads, it can deliver up to 1000 TOPS of NVFP4 compute. For FP32 workloads, it offers up to 31 TFLOPs of performance.
That makes DGX Spark particularly useful for running and experimenting with modern AI models. It can support local inference, AI agents, data science applications, visualization, and model development workflows that would normally require a much larger system.
The GPU also includes an additional 24 MB of L2 cache, helping improve performance in bandwidth-sensitive workloads.
128 GB unified memory for large AI models
One of the biggest advantages of DGX Spark is its 128 GB coherent unified system memory. Instead of separating CPU and GPU memory into different pools, the system uses a unified memory architecture. This allows the CPU and GPU to access a shared memory pool more efficiently.
The GB10 Superchip supports 256-bit LPDDR5x memory running at up to 9400 MT/s. This provides up to 301 GB/s of raw bandwidth, while the GPU can access up to 600 GB/s aggregate bandwidth over the high-speed C2C interface.
This memory configuration is one of the key reasons DGX Spark can work with large AI models. NVIDIA says the system can handle AI models with up to 200 billion parameters and fine-tune models with up to 70 billion parameters.
For developers working on large language models, this is a major advantage. Instead of relying completely on remote servers, DGX Spark allows more experimentation to happen locally, giving users greater control over privacy, cost, latency, and workflow speed.
4 TB NVMe storage and fast local access
The NVIDIA DGX Spark includes 4 TB of NVMe storage running at Gen5 x4 speeds. For AI workloads, fast storage is not just a convenience. Large models, datasets, checkpoints, and development environments can quickly consume space and require high read and write performance.
With 4 TB of fast local storage, DGX Spark gives users enough room to manage multiple models, datasets, containers, tools, and project files without immediately needing external storage.
ConnectX-7 networking and scalable AI performance
Another standout feature of DGX Spark is its networking capability. The system includes ConnectX-7 networking, allowing multiple DGX Spark units to be connected together.
By linking two systems, users can access a combined 256 GB of memory and work with larger AI models, including models with up to 405 billion parameters. The systems communicate through high-speed Ethernet using ConnectX technology, with the network interface connected through PCIe Gen5 x8.
This gives DGX Spark a level of flexibility that is rare in compact desktop systems. A single unit can function as a personal AI workstation, while multiple units can be combined into a small AI cluster. For teams experimenting with larger models, this can be an attractive alternative to jumping directly into large-scale infrastructure.
Compact design with a standard wall outlet
Despite its AI-focused hardware, DGX Spark remains compact enough to fit on a desk. It is designed to be power efficient and runs from a standard wall outlet. The GB10 Superchip has a TDP of 140W, which is modest considering the type of AI workloads the system is built to handle.
This matters because many AI development setups are bulky, loud, power-hungry, or complicated to deploy. DGX Spark takes a different approach by offering a clean desktop form factor that can still connect to external displays and peripherals like a normal workstation.
Display and connectivity support
DGX Spark supports up to four concurrent displays, including three DisplayPort outputs and one HDMI output. It can drive up to 4K at 120Hz using DisplayPort Alt Mode and up to 8K at 120Hz through HDMI 2.1a.
This makes the system suitable not only for AI development but also for multi-monitor productivity, visualization, and creative workloads. Users can monitor models, dashboards, terminals, notebooks, documentation, and visual outputs at the same time.
Security features are also built into the platform, including dual secure root support, SROOT and OSROOT processors, and support for both firmware TPM and discrete TPM.
DGX OS and software experience
The NVIDIA DGX Spark runs DGX OS, which is based on Ubuntu 24.04 LTS. For users coming from Windows, there is a learning curve. This system is not designed like a typical consumer PC where everything revolves around a familiar desktop app experience.
Instead, much of the work happens through the terminal, JupyterLab, the DGX Dashboard, and NVIDIA’s AI software tools. Developers who are already comfortable with Linux will feel more at home, while beginners may need time to adjust.
The system includes basic applications such as Firefox, a firmware updater, and standard Linux utilities. The interface is customized to match the DGX Spark experience. On first boot, the system may download updates, which can take some time depending on connection speed and update size.
DGX Dashboard and NVIDIA Sync
Newer DGX Spark units come with DGX Dashboard and NVIDIA Sync already installed. These two tools are important for managing and monitoring the system.
DGX Dashboard gives users a visual overview of system memory, GPU utilization, updates, and local applications. It also provides direct access to JupyterLab, which is where many AI development workflows take place.
NVIDIA Sync is especially useful when connecting multiple DGX Spark units. Instead of manually configuring everything through the terminal, Sync can simplify the process of linking systems into a cluster. In a dual-system setup, two DGX Spark units can run through a single IP over a 200 Gbps network, creating a compact but powerful local AI environment.
The Sync app can also test network speed and verify that the cluster is ready for local model workloads.
A personal AI cloud on your desk
One of the most appealing ways to think about DGX Spark is as a personal AI cloud. It can operate as a standalone AI workstation, but it can also be configured as a network-connected local AI system. This makes it useful for developers who want cloud-like flexibility while keeping workloads close to them.
Local AI development has several advantages. It can reduce dependency on external services, improve privacy for sensitive data, lower recurring cloud costs, and allow faster experimentation when working with models, agents, and prototypes.
For companies and research teams, DGX Spark can also serve as a bridge between local development and larger production environments. Workloads can be created and tested locally, then moved to DGX Cloud, accelerated data centers, or other cloud infrastructure when scaling becomes necessary.
Who is NVIDIA DGX Spark for?
The NVIDIA DGX Spark is not a typical Mini PC for everyday users. Its $4,699 price and specialized hardware make it a product for a specific audience.
It is best suited for AI developers, data scientists, machine learning engineers, researchers, robotics teams, software companies, university labs, and creators working with large AI models. It may also appeal to businesses that want a local system for experimenting with generative AI, private AI assistants, AI agents, and internal model development.
For casual users, traditional desktops or laptops will make more sense. But for people building with AI every day, DGX Spark offers something far more focused: a compact, efficient, and scalable AI development platform backed by NVIDIA’s software ecosystem.
Final thoughts
The NVIDIA DGX Spark 2025 represents a major step forward for local AI computing. It takes technologies normally associated with data center systems and brings them into a compact workstation that fits on a desk.
With the GB10 Grace Blackwell Superchip, 128 GB unified memory, 4 TB Gen5 NVMe storage, Blackwell GPU acceleration, ConnectX-7 networking, and support for large AI models, DGX Spark is built for the next generation of local AI development.
It is not the simplest machine for beginners, especially for users unfamiliar with Linux-based workflows. However, for developers and teams ready to work with local LLMs, AI agents, fine-tuning, and advanced AI pipelines, it offers a powerful and flexible foundation.
As AI workloads continue moving closer to users, systems like DGX Spark could become an important part of the future desktop landscape. It is small, efficient, highly specialized, and designed for one clear purpose: bringing serious AI supercomputing power to the personal workspace.NVIDIA DGX Spark Review: A Compact AI Supercomputer Built for Local LLMs, Developers, and Multi-Node Workloads
The NVIDIA DGX Spark is not just another small workstation. It is a compact AI system designed for developers, researchers, and creators who want serious local AI performance without relying on cloud infrastructure or a full data-center setup. As AI models continue to grow larger and more demanding, the appeal of a quiet, desk-friendly machine capable of running massive language models locally becomes much easier to understand.
At its core, DGX Spark is built around NVIDIA’s GB10 Superchip, combining a 20-core Arm CPU, powerful Blackwell-generation GPU capabilities, 128 GB of coherent unified memory, fifth-generation Tensor Cores, NVLink-C2C, and high-speed networking support through ConnectX-7. The result is a tiny system that can handle workloads normally associated with much larger machines.
For many local AI developers, the real sweet spot may be two or four DGX Spark systems working together. A single unit is already capable, but clustering multiple systems opens the door to larger models, more VRAM, and much stronger inference performance.
In CPU testing, DGX Spark delivers respectable performance for its size and power target. Running Geekbench 7, the system scored 2,568 points in single-core and 23,656 points in multi-core performance.
The single-core result is solid, especially considering the Arm-based 20-core CPU is not chasing extremely high clock speeds. It even manages to outperform AMD’s Ryzen AI MAX+ 395 in this specific test, despite that chip reaching much higher boost frequencies. DGX Spark is tuned more for efficient, continuous workstation use than for short bursts of peak CPU performance.
Multi-core performance is where the system becomes more impressive. DGX Spark comes surprisingly close to higher-thread-count x86 desktop processors while staying ahead of several mobile SoCs. This is important because AI development work often involves more than GPU inference. Dataset preparation, model loading, prompt handling, software compilation, and local services can all benefit from strong CPU throughput.
Future RTX Spark PCs are expected to use similar CPU technology with more optimized clocks, so they may deliver better results in traditional CPU benchmarks. DGX Spark, however, is clearly built around 24/7 reliability and local AI development rather than being a conventional desktop replacement.
Although DGX Spark is not marketed as a gaming PC, its GPU hardware gives a glimpse of what upcoming RTX Spark systems may offer. To test this, Cyberpunk 2077 was run through Steam on Linux using Proton Experimental and the Steam Linux Runtime. After setup, the game ran properly with RTX features enabled.
The GB10 GPU has a core count comparable to an RTX 5070, but the power budget is much tighter. DGX Spark is rated around 140W total, and that budget is shared across the GPU, CPU, and memory. In practice, the GPU likely operates closer to the 90W to 100W range under gaming loads.
At 1440p using the RT Ultra preset, Cyberpunk 2077 averaged 28.70 FPS. At 1080p with the same ray tracing preset, it reached 59.34 FPS. With high settings at 1080p and ray tracing disabled, performance exceeded 100 FPS, which is a strong result for a system primarily designed for AI workloads.
The bigger difference comes from NVIDIA’s latest AI-assisted rendering technologies. With DLSS Multi Frame Generation, DLSS 4.5 Transformer, and DLSS 4.5 Ray Reconstruction enabled, Cyberpunk 2077 reached up to 95.15 FPS at 1440p with RT Ultra and around 140 FPS at 1080p.
That does not make DGX Spark a gaming-first machine, but it does show how much potential the RTX Spark family could have once the hardware is paired with better gaming-focused optimization and Windows support.
The main reason to consider DGX Spark is local AI performance, especially for large language models. This system is aimed at developers who want to run, test, fine-tune, and experiment with modern LLMs without depending entirely on cloud services.
Testing covered several popular and newer AI models using llama.cpp, chosen for its low overhead and efficient local inference behavior. NVIDIA also provides setup playbooks for multiple inference platforms, including llama.cpp, vLLM, LM Studio, and other tools, making it easier to get started with common AI development workflows.
Across models ranging from 4 billion parameters up to 120 billion parameters, DGX Spark delivered strong token generation and prompt processing performance. Its compact size makes the results even more impressive. This is not a large tower workstation packed with multiple desktop GPUs. It is a small, quiet system that can sit on a desk and still run heavyweight local AI workloads.
Mixture of Experts models were also tested, including single-node and dual-node configurations. These MoE models are designed to offer the knowledge capacity of very large models while activating only part of the model during inference, improving efficiency.
DeepSeek V4 Flash with NVFP4 delivered 21.3 tokens per second on a single DGX Spark. When two DGX Spark systems were clustered together, performance increased to 59.13 tokens per second. That is more than a 2x improvement and highlights the advantage of pairing two systems to access 256 GB of combined memory for very large models.
Some models simply cannot fit comfortably on a single unit. MiniMax-M2.5 in 4-bit form could not run within one Spark, but a dual-system cluster handled it at 40.12 tokens per second. NVIDIA’s open-source Nemotron 3 Super with 120 billion parameters also performed well on the clustered setup.
This is where DGX Spark begins to feel less like a compact PC and more like a personal AI cluster. For developers working with agents, retrieval-augmented generation, coding models, synthetic data, and large-context experiments, the ability to scale from one box to two with minimal physical space and power draw is a major advantage.
Software support is just as important as hardware for this kind of product. NVIDIA’s software stack includes DGX OS, a system dashboard, NVIDIA Sync, and a growing collection of guided setup playbooks. These tools help turn what could have been a complex Linux-only experiment into a more usable AI development environment.
That said, the setup experience is not always perfect. Some guides may require extra troubleshooting, and not every playbook works as a simple copy-and-paste process. Depending on the model, inference engine, and repository, users may still need to make adjustments to get the best performance. Experienced developers will likely be comfortable with that, but newcomers should expect some hands-on configuration.
NVIDIA is continuing to improve the platform with updates focused on faster inference and easier deployment of agentic AI tools. One-click installation options for AI agent platforms are expected to make DGX and RTX Spark systems more approachable for developers building local AI assistants, coding agents, research tools, and automated workflows.
Generative AI image workloads were also tested using ComfyUI with models such as Flux.2 and SDXL Turbo. Setup was relatively straightforward, and the system produced around six to seven 1024 x 1024 images per minute. That makes DGX Spark useful not only for language models but also for local image generation, creative experimentation, and AI-assisted content workflows.
Thermals are reasonable for a compact system with this level of performance. DGX Spark does get warm under heavy workloads, but fan noise remains low. Exhaust heat is noticeable when the system is pushed hard, yet the machine does not become disruptive in a typical workspace.
Idle temperatures were around 48°C, while peak temperatures reached approximately 81°C. That is warm, but not unusual for a dense workstation designed to sustain AI workloads in a small chassis.
Power consumption is one of DGX Spark’s most attractive qualities. Although the system is rated up to 140W, real-world testing rarely showed power draw exceeding 100W, even in demanding prefill-heavy scenarios. A single unit typically averaged around 80W, while a second paired unit often consumed slightly less than the primary system.
That efficiency matters. It means developers can run one or two systems for long experiments, overnight inference jobs, or local agent testing without dealing with excessive noise, heat, or electricity costs. For a personal AI workstation, that balance of performance and efficiency is one of the strongest selling points.
The DGX Spark ultimately succeeds because it defines a new category: the compact local AI supercomputer. It is not the cheapest way to run AI models locally, with a single unit priced around $4,699, but it is one of the most complete and scalable options available in a small form factor.
Its biggest strengths are clear: 128 GB of unified memory per system, strong LLM inference performance, efficient power use, quiet operation, multi-node scalability, and a software ecosystem designed specifically for AI developers. A single DGX Spark can handle serious local AI tasks, while two units become a powerful 256 GB cluster capable of running models that would otherwise require much larger systems.
Gaming is only a side benefit, but the results show the hardware has plenty of graphics potential. The real purpose of DGX Spark is local AI development, including large language models, AI agents, RAG pipelines, generative AI, and multi-node experimentation.
A year after its introduction, DGX Spark still feels like the reference point for compact AI workstations. It proves that a personal AI supercomputer can sit quietly on a desk, run from a standard wall outlet, and handle workloads that once felt out of reach without cloud compute or enterprise hardware.
For developers who need serious local AI capability in a compact and efficient system, NVIDIA DGX Spark remains one of the most compelling options available.






