NVIDIA Vera Rubin Promises a Major Leap in Agentic AI Inference Performance and Token Efficiency
NVIDIA’s next-generation Vera Rubin platform is shaping up to be a major breakthrough for agentic AI workloads, delivering far higher token throughput and dramatically lower operating costs compared with current Blackwell systems. While Blackwell already represents a huge jump over Hopper for AI inference, Vera Rubin appears to move into an entirely different performance class, especially for AI agents, coding assistants, and large-scale inference deployments.
The latest performance figures are based on on-silicon testing using real-world agentic coding trajectories. These workloads are designed to reflect how AI agents actually behave when handling coding tasks, including long-context reasoning, repeated tool use, and multi-step generation. The benchmark used in the testing evaluates AI infrastructure across advanced models such as Kimi K3, MiniMax M3, GLM5.3, Qwen3.5, and DeepSeek V4 Pro.
A key focus of the testing is not just raw speed, but usable performance. For agentic AI, responsiveness matters as much as total throughput. Metrics such as end-to-end interactivity, standard interactivity, end-to-end latency, and time to first token help show how quickly users receive responses and how efficiently the system delivers tokens under real workloads.
End-to-end interactivity measures the average output-token rate per user across the full request, from submission to final token delivery. This is important because it includes the waiting time before the first response appears. Standard interactivity focuses only on the generation phase, measuring the token rate from the first token to the last. End-to-end latency tracks the total time needed to complete a request, while time to first token shows how quickly the system begins responding. For AI agents that repeatedly start new long-context turns, a fast first response is critical.
Blackwell already delivers a major generational improvement. NVIDIA’s GB300 NVL72 Grace Blackwell server is said to provide up to 15 times higher throughput per megawatt than H200 NVL8 Hopper systems when running DeepSeek V4 Pro 1.6T. That means Blackwell can support much more responsive agentic inference within the same power envelope, which is increasingly important as AI data centers face strict power and infrastructure limits.
Cost efficiency is another major improvement. Blackwell reportedly delivers up to 10 times lower cost per million tokens than Hopper. For AI operators, this can translate into more agent capacity using the same infrastructure budget, or similar capacity at a much lower operating cost. As demand for AI coding agents, enterprise assistants, and automated reasoning tools continues to rise, reducing token cost is becoming one of the most important factors in AI infrastructure planning.
Blackwell also appears to scale well with larger models. In testing with Kimi K3 2.8T, the Grace Blackwell platform delivered up to 80 times higher throughput per megawatt compared with Hopper. It also sustained around 215 tokens per second per user in interactivity, a level that significantly exceeds what Hopper-class systems can deliver. This suggests Blackwell is not only faster, but also better suited for increasingly large and complex AI models.
However, Vera Rubin is where NVIDIA’s agentic AI roadmap becomes even more aggressive.
The NVIDIA Vera Rubin NVL72 platform is claimed to deliver a massive 30 times increase in throughput compared with Grace Blackwell NVL72 in DeepSeek V4 Pro 1.6T workloads. In terms of interactivity, Vera Rubin reaches around 160 tokens per second per user and can scale up to roughly 280 tokens per second per user. By comparison, Blackwell peaks below 180 tokens per second per user in the same class of testing.
That kind of gain could be highly significant for AI factories, cloud providers, and enterprises building large-scale agentic AI services. Higher token throughput means more simultaneous users, more active agents, faster coding workflows, and better responsiveness in complex reasoning tasks. For applications such as AI software development, automated debugging, workflow orchestration, and enterprise copilots, these improvements could directly impact user experience and service capacity.
Vera Rubin also brings a major reduction in token cost. In agentic coding workloads, the platform is said to offer up to 35 times lower cost per million tokens compared with Blackwell. This is one of the most important figures for companies planning to run AI agents continuously at scale. Lower token cost can make always-on agents more practical, allowing businesses to deploy larger numbers of AI workers without overwhelming power and budget constraints.
NVIDIA is also using power-management technologies to improve deployment efficiency. With systems designed to manage power across GPUs, racks, and workloads, AI factories may be able to provision up to 40% more GPUs within the same megawatt budget. This matters because power availability has become one of the biggest bottlenecks in modern AI infrastructure. Even if companies can buy more accelerators, they still need enough energy and cooling capacity to run them efficiently.
Another important detail is that the reported Vera Rubin results do not yet fully reflect the performance contribution of the Vera CPU for tool calling. The current figures focus mainly on the Vera Rubin chips rather than the complete multi-chip platform design. That means future results could show even stronger performance once the full platform is deployed and optimized across CPUs, GPUs, networking, memory, and software.
This is especially relevant for agentic AI, where tool calling can play a major role. AI agents often need to interact with external tools, execute code, retrieve data, analyze files, and perform multi-step operations. CPU performance, memory bandwidth, networking, and system-level optimization all influence how smoothly these workflows run. As Vera Rubin systems mature, the full platform could deliver additional gains beyond raw GPU inference throughput.
NVIDIA’s broader strategy is clear: keep improving performance not only through new silicon, but also through software stack optimizations, platform-level design, and power-aware infrastructure management. Blackwell has already shown how much optimization can improve agentic inference performance, but Vera Rubin appears positioned to push AI throughput and cost efficiency much further.
For the AI industry, the implications are substantial. If Vera Rubin delivers the claimed performance at scale, it could lower the cost of running advanced AI agents, accelerate the adoption of autonomous coding systems, and make large-model inference more accessible for high-demand enterprise and cloud workloads. Faster tokens, lower latency, and reduced cost per million tokens are exactly what AI service providers need as user demand continues to grow.
In short, Blackwell is a major leap over Hopper, but Vera Rubin looks like NVIDIA’s next big step toward high-efficiency, large-scale agentic AI. With up to 30 times higher throughput than Grace Blackwell in key workloads, up to 35 times lower token cost, and improved power efficiency for AI factories, Vera Rubin could become one of the most important platforms for the next generation of AI inference.






