Nvidia’s eagerly anticipated Blackwell GPUs, which have already experienced delays, are now under scrutiny due to significant overheating challenges in high-capacity server racks. This development threatens to further postpone their release, a concern that resonates across the tech industry as AI accelerator shortages continue to impact data centers worldwide.
The problematic overheating centers around Nvidia’s NVL72 racks, which can hold as many as 72 Blackwell chips, purposed for AI and high-performance computing. These racks demand a staggering 120 kW per unit, which makes them susceptible to performance degradation and even hardware damage if not managed correctly. Weighing in at 1.5 tons and consuming up to 132 kW, these units have set a record for single-server power consumption, a feat that, while impressive, also raises questions about their viability under current design structures.
The issue has rung alarms for customers eager to deploy next-generation AI data centers, prompting Nvidia to press its suppliers for revised rack designs. Despite the mounting concerns, Nvidia has remained somewhat nonchalant, with a spokesperson stating, “The engineering iterations are normal and expected,” suggesting collaborations with leading cloud service providers are ongoing.
Originally slated for a late 2024 release following its March 2024 announcement, Blackwell promised groundbreaking performance improvements over its predecessors. However, due to these design complications, its launch was deferred to early 2025. This delay poses potential challenges for industry giants like Meta, Google, and Microsoft, who rely heavily on Nvidia’s AI hardware. The delays might also challenge Nvidia’s stronghold on market leadership unless rectifications are promptly and effectively made.
From a financial perspective, the stakes are high. Each unit of the GB200 Grace Blackwell superchip demands an eye-watering $70,000, with a full server rack surpassing $3 million. Nvidia has ambitious sales targets, aiming to move between 60,000 and 70,000 servers, making any further deferrals both costly and strategically critical.
As the market watches how Nvidia navigates these turbulent waters, Wall Street firms are cautiously adjusting their projections. Morgan Stanley’s initial high hopes for shipments in late 2024 have been dampened, while Raymond James remains somewhat optimistic for the long-term demand in 2025. Meanwhile, Jefferies warns of potential overestimation of future revenue and profits, estimating that Nvidia’s shipments could be lower than anticipated.
Despite the intensified scrutiny, analysts like Holger Mueller of Constellation Research emphasize the integral role cooling systems play in the sustainable operation of AI platforms. Mueller asserts that cooling inadequacies could directly impact Nvidia’s prominent position if the issue isn’t resolved swiftly. How Nvidia manages this crisis and the timing of their resolution could influence their stock price significantly, an element that both investors and competitors are closely monitoring.






