A close-up of a server rack filled with rows of NVIDIA processing units, illuminated under stage lighting.

NVIDIA’s Rubin Ultra Rack Could Hit $21 Million as HBM4e Memory Costs Soar to $1.5 Million Per Unit

NVIDIA Rubin and Rubin Ultra AI Servers Could Push Data Center Costs to Record Highs

NVIDIA’s next generation of AI servers is shaping up to be far more expensive than the already costly Blackwell platform. New estimates suggest that Rubin and Rubin Ultra systems will bring a major jump in rack-level pricing, with high-bandwidth memory playing one of the biggest roles in the increase.

The key reason is simple: next-generation AI performance demands enormous amounts of memory. Rubin systems are expected to use HBM4, while Rubin Ultra will move to HBM4e. These memory technologies are faster and more advanced than the HBM3e used in Blackwell, but they also come with a much higher price tag. As AI models continue to grow in size and complexity, memory capacity and bandwidth are becoming just as important as raw GPU compute power.

According to industry estimates, NVIDIA’s Rubin VR200 platform, also known as Oberon, will use NVL72 racks equipped with 72 Rubin GPUs and 36 Vera CPUs. Each Vera Rubin tray is expected to include four Rubin GPUs and two Vera CPUs. A smaller motherboard unit, referred to as a Superchip, will contain two GPUs and one CPU. With 36 Superchips per NVL72 rack, the full system reaches 72 GPUs and 36 CPUs.

Each Rubin GPU is expected to feature 288 GB of HBM4 memory. That gives a full NVL72 rack a total of 20.7 TB of HBM4 memory. On top of that, each Vera CPU is expected to include 1.5 TB of LPDDR5X memory, bringing the total LPDDR5X capacity per rack to 54 TB. These are massive memory figures, even by modern AI data center standards.

But memory is only part of the complete server cost. A Rubin rack also requires advanced networking, cooling, power delivery, interconnects, CPUs, motherboards, and other infrastructure. Still, HBM4 stands out as one of the most expensive components in the system.

For comparison, NVIDIA’s Blackwell B200 NVL72 rack uses 72 GPUs, with each GPU carrying 192 GB of HBM3e memory. That gives the rack 13,824 GB of HBM3e. The estimated cost of HBM3e for a Blackwell B200 rack is around $156,000, with memory priced at roughly $11.26 per GB.

Rubin raises those numbers sharply. A Rubin V200 NVL72 rack is expected to include 20,736 GB of HBM4 memory, with the cost per GB estimated at around $18.40. That pushes the total HBM4 cost per rack to about $382,000. In other words, the HBM cost alone is expected to be more than twice as high as Blackwell.

The total average selling price of a Blackwell B200 rack is estimated at around $3 million. For Rubin V200, that figure could rise to around $6 million per rack. HBM would account for about 6.4 percent of the total rack cost, compared with 5.2 percent for Blackwell B200.

Rubin Ultra takes the price increase even further.

The Rubin Ultra V300 platform, also known as Kyber, is expected to use NVL144 racks with 144 GPUs. Each V300 GPU is expected to carry 576 GB of HBM4e memory, doubling the memory capacity per GPU compared with standard Rubin V200.

That would give a full Rubin Ultra NVL144 rack an enormous 82,944 GB of HBM4e memory. Even if the cost per GB remains close to standard Rubin at around $18.49, the total HBM cost per rack would climb to roughly $1.534 million.

That is about four times the estimated HBM cost of a standard Rubin rack and nearly five times higher than Blackwell Ultra. Blackwell Ultra B300 racks are expected to include 72 GPUs, each with 288 GB of HBM3e, for a total of 20,736 GB of memory per rack. The estimated HBM cost for Blackwell Ultra is around $317,000.

The full rack price difference is even more dramatic. Blackwell Ultra is estimated at around $4 million per rack, while Rubin Ultra could reach approximately $21 million per rack. Although HBM would represent about 7.3 percent of the Rubin Ultra rack cost, the absolute memory bill is still enormous.

These figures show how aggressively AI infrastructure costs are rising as the industry moves into the Rubin generation. NVIDIA’s future AI servers are not simply adding more GPUs; they are scaling up memory capacity, bandwidth, rack density, and system complexity at the same time.

The cost impact also extends beyond the hardware purchase price. Large-scale Vera Rubin AI data centers are expected to require massive power and cooling investments. Estimates suggest that a one-gigawatt AI data center based on Vera Rubin systems could cost more than $47 billion to develop. Annual power expenses alone could reach around $1.3 billion.

That means Rubin will be expensive to buy, expensive to deploy, and expensive to operate. However, NVIDIA is positioning the platform around total cost of ownership rather than only upfront cost. The company’s argument is that higher performance, better efficiency, and faster AI processing can reduce the cost per generated token, which is one of the most important metrics for large AI service providers.

For companies training and running massive AI models, the economics are not just about server pricing. They also care about throughput, energy efficiency, model response speed, data center density, and long-term operating cost. If Rubin delivers a major performance leap over Blackwell, the higher rack price may still be attractive to hyperscalers and AI firms that need maximum compute capacity.

Demand for Rubin is expected to be extremely strong. Blackwell already triggered a major wave of AI infrastructure investment, and Rubin could push that momentum even further. As AI companies race to build larger models and serve more users, access to advanced GPU clusters remains a strategic priority.

The move from Blackwell to Rubin also highlights the growing importance of memory in AI computing. In earlier GPU generations, much of the attention focused on compute cores and raw processing power. Today, memory bandwidth and capacity are just as critical. HBM4 and HBM4e will help feed data to the GPUs faster, enabling better performance for training, inference, and large-scale generative AI workloads.

The downside is clear: premium memory technology is becoming one of the biggest cost drivers in next-generation AI servers. With Rubin and Rubin Ultra, NVIDIA’s rack prices could reach levels never seen before in mainstream data center hardware.

Still, the AI industry appears ready to absorb these costs. For the largest cloud providers, AI labs, and enterprise infrastructure buyers, the question is not whether these systems are expensive. The real question is whether they can deliver enough performance and efficiency to justify the investment.

If NVIDIA’s Rubin platform meets expectations, it could become the next major foundation for global AI infrastructure. But one thing is already clear: the next era of AI computing will be faster, denser, and significantly more expensive.