A compact, unbranded server unit is placed on a pedestal in a dimly lit data center corridor lined with racks of server equipment.

Home AI Powerhouse Runs 552B DeepSeek V4.1-Flash at 494 Tokens/Second on 4 NVIDIA DGX Spark Units

A powerful at-home AI setup built from four NVIDIA DGX Spark systems has shown just how fast local inference can be when the right hardware is paired with an efficient open-weight model. Running DeepSeek V4.1-Flash, the compact four-node cluster reportedly reached a peak output of nearly 500 tokens per second, putting serious local AI performance within reach for users who want speed, privacy, and control without relying on cloud APIs.

For companies, researchers, and developers handling sensitive data, local AI hosting remains one of the most secure options. Instead of sending prompts, code, documents, or proprietary business information to external services, a self-hosted AI model keeps everything on private hardware. The tradeoff is usually cost and setup complexity, but this latest configuration shows that a relatively compact system can deliver impressive throughput even with a very large model.

The setup was shared by tech analyst Patrick Moorhead, who revealed an at-home AI rack made up of four NVIDIA DGX Spark units. Together, the systems provide around 512 GB of pooled unified memory. The machines were stacked in a compact rack with additional cooling fans underneath, creating a dense local AI workstation designed for high-performance inference.

Each NVIDIA DGX Spark system includes a 20-core Arm CPU with 10 high-performance Cortex-X925 cores and 10 efficiency-focused Cortex-A725 cores. The platform also includes NVIDIA’s GB10 integrated GPU, which features fifth-generation Tensor Cores, RTX ray tracing cores, DLSS 4 support, up to 31 TFLOPs of FP32 performance, and up to 1000 TOPS of NVFP4 compute for AI workloads.

A key advantage of DGX Spark is its 128 GB coherent unified LPDDR5X memory, shared between the CPU and GPU. This memory delivers around 273 GB/s of bandwidth over a 256-bit interface, which is especially useful for running large language models locally. The system also includes high-speed networking through ConnectX-7 SmartNIC hardware with dual 200 Gb/s QSFP ports, along with 10 GbE, Wi-Fi 7, Bluetooth, HDMI 2.1a, USB-C connectivity, and NVMe SSD storage options typically ranging from 1 TB to 4 TB.

Instead of using a traditional Ethernet switch, the four DGX Spark units in this setup were connected through a switchless ring topology using QSFP cables. In this layout, the first system connects to the second and fourth units, the second connects to the first and third, the third connects to the second and fourth, and the fourth connects back to the third and first. NVIDIA generally recommends a 200 GbE switch in a star topology for four DGX Spark units, but skipping the switch helps reduce the total system cost.

The model chosen for this local AI cluster was DeepSeek V4.1-Flash, a 552-billion-parameter model designed for high efficiency. Despite its massive total parameter count, the model activates only a small portion of its parameters during inference. It uses 8 billion active parameters during the prefill stage and 16 billion during the decode stage. The architecture includes one shared expert and 384 routed experts, with six routed experts active at a time.

DeepSeek V4.1-Flash also uses several efficiency-focused design features, including Causal Encoder-Decoder architecture, Compressed Sparse Attention 2, FP4 quantization, and SWA Bounded Replay. These optimizations help reduce the KV cache requirement to only 890 bytes per token, which makes the model far more practical to run on advanced local hardware.

The performance numbers are the most eye-catching part of the setup. In code generation workloads with 32 concurrent requests, the four-node NVIDIA DGX Spark cluster reportedly reached a peak output of 494 tokens per second. For prose generation, the system peaked at 280 tokens per second with the same level of concurrency. In single-request testing, the setup delivered around 58 tokens per second for prose and roughly 96 tokens per second for code.

Prompt processing was also strong, reaching approximately 4,764 tokens per second. First-token latency while idle was around 0.2 seconds, making the system feel highly responsive. The total memory footprint for the model was listed at 476 GB, which fits within the combined 512 GB memory pool of the four DGX Spark units.

The cost, however, is not small. A single 128 GB NVIDIA DGX Spark unit can currently range from about $5,500 to more than $9,000 depending on storage capacity and availability. Four systems would therefore cost roughly $22,000 to $36,000 before adding accessories. QSFP direct-attach cables for the ring configuration may add a few hundred dollars, while the rack, cooling fans, power distribution gear, and other components can add several hundred more.

Altogether, the full local AI rig would likely cost somewhere between $25,000 and $40,000. That puts it well outside the range of casual users, but it could make sense for AI developers, research teams, small companies, and professionals who frequently run large models and want to avoid ongoing API fees, rate limits, or data privacy concerns.

While cloud-based AI subscriptions remain easier to use, local inference offers major advantages for certain workloads. A privately hosted model can run continuously, process confidential data, avoid third-party usage restrictions, and provide predictable performance without depending on external servers. For organizations working with sensitive code, legal files, financial data, medical documents, or internal research, that level of control can be extremely valuable.

This four-node NVIDIA DGX Spark and DeepSeek V4.1-Flash setup highlights a growing trend in AI computing: powerful local systems are becoming capable enough to challenge cloud inference for specialized users. The hardware is still expensive, but the combination of unified memory, high-speed interconnects, and efficient model architecture shows how much performance can now be packed into a compact home or office AI rack.

For anyone watching the future of local AI, this result is a strong sign of where the industry is headed. Large language models are becoming more efficient, compact AI hardware is becoming more capable, and self-hosted inference is moving from an enthusiast experiment to a realistic option for serious professional use.