A large, vertically standing, unbranded server rack filled with equipment is being installed by workers in a dimly lit data center.

NVIDIA Scales Back Vera Rubin Memory as Soaring HBM4 Costs Pressure Rack Economics

Nvidia Reportedly Cuts Vera Rubin NVL72 Memory to Control AI Rack Costs Amid Supply Crunch

Nvidia is reportedly making major changes to the memory configuration of its upcoming Vera Rubin NVL72 rack-scale AI system as the company looks to manage rising component costs and ongoing supply pressure in the memory market.

According to analysis from GF Securities, Nvidia plans to reduce the amount of SOCAMM memory used in the Vera Rubin NVL72 platform, a move that could significantly lower the total cost of the system. The adjustment comes as demand for advanced memory continues to surge across the AI industry, pushing prices higher and tightening availability for key components such as LPDDR5X and HBM.

The Vera Rubin NVL72 is expected to be one of Nvidia’s most powerful AI systems, designed for large-scale data centers running demanding artificial intelligence workloads. Earlier performance figures suggested a dramatic leap over current-generation Blackwell-based systems. While Blackwell AI platforms were reported to deliver around 80,000 tokens per second for mixture-of-experts workloads at 150 megawatts of power consumption, the VR200 Vera Rubin NVL72 platform reportedly reaches around 800,000 tokens per second, representing a tenfold improvement.

However, that level of performance comes with a steep price tag. Previous estimates suggested that a single Rubin NVL72 rack could cost as much as $9.1 million, largely because of elevated memory pricing. Earlier lower estimates, reportedly around $7.8 million per rack, may no longer reflect current market conditions as memory prices continue to rise. HBM4 memory, which will be a critical part of next-generation AI systems, is expected by some analysts to become considerably more expensive in the coming years.

GF Securities now believes Nvidia is responding by reducing the default SOCAMM configuration in Vera Rubin NVL72 racks. Instead of using 192GB SOCAMM modules, Nvidia is reportedly moving to 96GB modules, effectively cutting that portion of memory capacity in half. This change is believed to be linked to constraints in the LPDDR5X market, where supply remains under pressure due to strong demand from AI, mobile, and high-performance computing customers.

The report also indicates that memory capacity on Vera CPUs could be reduced to around 28TB, down from earlier expectations of roughly 54TB to 55TB. The GPU side of the system, however, is expected to retain 20.7TB of HBM4 memory per rack. For CPU racks, Nvidia is also expected to adopt 96GB SOCAMM per CPU.

The financial impact could be substantial. GF Securities estimates that reducing LPDDR5X capacity could bring memory costs for the VR200 system down sharply. If capacity is reduced to one-fourth of the original level, LPDDR5X costs could fall to around $293,000. If capacity is cut in half, costs could decline to about $586,000, instead of reaching the earlier estimate of approximately $1.2 million.

Without these adjustments, memory costs could account for roughly 29% of the total bill of materials for the VR200 platform, or about $2.1 million. That figure is well above the preferred level of around 20%, making memory cost management a key priority for Nvidia as it prepares its next wave of AI infrastructure.

The move highlights a growing challenge across the AI hardware industry. As companies race to build larger and faster AI clusters, memory has become one of the most important and expensive parts of the system. High-bandwidth memory is essential for training and inference workloads, while LPDDR5X and SOCAMM configurations help support massive CPU and system-level memory requirements.

Nvidia appears to be using its supply chain strength to navigate these pressures. The company reportedly secured long-term memory agreements before the shortage became more severe, helping it avoid some of the disruption faced by other technology firms. Even so, the latest reported changes show that no company is fully insulated from rising memory prices and limited supply.

For data center customers, the reduced memory configuration could make the Vera Rubin NVL72 more financially viable while preserving its most important AI performance advantages. For Nvidia, the adjustment may help protect margins and keep its next-generation AI rack systems competitive as infrastructure costs continue climbing.

The Vera Rubin NVL72 remains one of the most anticipated AI platforms on the roadmap, and its combination of massive compute power, HBM4 memory, and rack-scale design is expected to play a major role in future AI data centers. But as this report suggests, the next phase of AI hardware competition will not be defined by performance alone. Memory cost, supply availability, and system-level efficiency may become just as important.