NVIDIA’s Blackwell Ultra GB300 AI Racks Crush Long-Context DeepSeek Workloads, Leaving GB200 Behind

NVIDIA’s GB300 NVL72 AI racks are already showing strong real-world momentum after being tested across DeepSeek’s newest open-source models. With a mix of fine-tuning and inference-side optimizations, the early numbers point to meaningful gains for organizations running long-context AI workloads, especially those that care about responsiveness and real-time output.

The big story with GB300, built around NVIDIA’s Blackwell Ultra platform, is its push toward better long-context performance. That matters because today’s agentic AI systems don’t just answer short prompts—they keep track of lengthy conversations, large documents, tool outputs, and evolving memory. As context windows grow, the workload pressure increasingly shifts toward GPU memory capacity and how efficiently the system can move through prompt processing and token generation without choking on bottlenecks.

To evaluate long-context inference performance, the Large Model Systems Organization (LMSYS) tested GB300 NVL72 and incorporated infrastructure-level routing techniques intended for large-scale deployments. One key method used was PD (Prefill-Decode) Disaggregation, a practical approach for scaling large-context serving. In straightforward terms, the “prefill” stage handles prompt processing, while the “decode” stage generates tokens. By splitting these phases across different nodes, systems can reduce bottlenecks and improve overall throughput—particularly when many users are running long-context requests at once.

Alongside disaggregation, the testing also included optimization strategies such as dynamic chunking to improve responsiveness under long context windows, plus more effective handling and translation of KV (key-value) cache capacity—an important factor in keeping long-context inference efficient.

In the reported comparison between NVIDIA’s GB300 NVL72 and GB200 NVL72, the highlights were clear:
– Up to 1.53x peak throughput, reaching 226.2 tokens per second per GPU (TPS/GPU)
– Up to 1.87x user speed, with a major jump in tokens per second per user attributed to MTP (Multi-Token Prediction)
– Up to 1.58x improved latency

LMSYS notes that GB300 typically delivers around a 1.4x to 1.5x advantage over GB200, with the strongest gains showing up in latency-sensitive scenarios. That emphasis on lower latency fits directly with where the market is heading: agentic AI tools that need to “think” and respond quickly, handle long inputs, and take actions across multiple steps without feeling sluggish.

While these performance and latency improvements make Blackwell Ultra look like a standout option for hyperscalers and newer cloud providers, total cost of ownership remains an open question in the broader industry conversation. GB300 deployments can come with higher costs, and many buyers will want more clarity on the full price-to-performance equation at scale.

Even so, the direction is consistent with NVIDIA’s recent generational approach: not only boosting raw architecture performance, but also targeting the practical constraints that define modern AI serving—long context, high concurrency, and the need for fast, dependable responses. For organizations betting on long-context inference and agentic workflows, GB300 NVL72 is quickly shaping up as one of the most compelling platforms in the current AI infrastructure race.