NVIDIA Vera Rubin Pushes AI Data Centers Toward Higher Token Throughput and Better Power Efficiency
NVIDIA is putting energy efficiency at the center of its next-generation AI infrastructure strategy, and its Vera Rubin platform is shaping up to be a major part of that vision. As artificial intelligence workloads continue to grow, data centers are under increasing pressure to deliver more performance without demanding more electricity from already stressed power grids.
At the AI Infra Summit, NVIDIA highlighted how its full-stack AI factory approach is designed to improve token throughput, reduce wasted power, and help operators get more value from every megawatt. Instead of focusing only on raw compute performance, NVIDIA is pushing a new metric for AI infrastructure: tokens per megawatt.
That shift matters. Modern AI systems are no longer judged only by how many floating-point operations they can perform. For inference, agentic AI, copilots, search, and enterprise automation, what really matters is how many useful tokens can be generated reliably, quickly, and affordably within a fixed power envelope.
NVIDIA’s AI factory platform combines Vera Rubin systems, Dynamo inference software, NeMo libraries, and high-speed networking technologies such as NVLink, Spectrum-X Ethernet, ConnectX SuperNICs, BlueField DPUs, and BlueField-powered infrastructure services. Together, these components are meant to help AI data centers scale to thousands of nodes while keeping power, networking, security, and storage efficiency under control.
One of the key technologies NVIDIA discussed is DSX MaxLPS, which is designed to raise token throughput per megawatt by as much as 40%. For AI data center operators, that kind of improvement can be extremely valuable because expanding power capacity is often expensive, slow, or impossible in the short term.
In practice, this means operators may be able to serve more AI users, run more inference jobs, and increase overall cluster productivity without adding new power lines or expanding the physical footprint of a facility.
NVIDIA also pointed to DSX Flex, a system that allows AI data centers to respond dynamically to power grid conditions. When the grid is under stress, the platform can adjust power demand without interrupting important AI workloads. This kind of grid-aware orchestration is becoming increasingly important as AI facilities consume more electricity and utilities look for ways to maintain stability during peak demand.
Emerald AI has been using DSX through its Conductor platform, demonstrating how AI workloads can be coordinated with real-world grid conditions. According to the details shared, the system has already responded to around 200 grid-condition signals and successfully adjusted operations each time without user intervention.
The idea is simple but powerful: instead of treating an AI data center as a fixed power consumer, DSX Flex allows it to behave more intelligently. It can reduce demand when the grid is constrained while still protecting priority workloads and maintaining service reliability.
Lambda, a GPU cloud provider serving more than 10,000 customers, tested the software on a five-rack, 19-node cluster. The result was a notable improvement in efficiency. By operating 19 nodes within the same power budget normally used by 16 nodes at full power, the company achieved around 24% more cluster-wide token throughput.
Throughput rose from roughly 4 million tokens per second to about 5 million tokens per second, while performance per watt improved by 23%. For cloud providers, this type of gain directly affects infrastructure economics, allowing them to deliver more AI compute capacity without proportionally increasing power costs.
NVIDIA says these improvements can translate into major gains across Vera Rubin and Groq 3 LPX racks. In some configurations, operators could fit up to 40% more GPUs within the same site power envelope, while achieving up to 35% higher token throughput without requiring new power lines.
That is especially important for AI companies and cloud providers facing power limitations. In many markets, power availability has become one of the biggest bottlenecks for building new AI capacity. If software and system-level optimization can increase output within existing limits, operators can delay or avoid expensive infrastructure upgrades.
The performance numbers also show how these systems could benefit demanding AI workloads. On a 100K-context Qwen 3.8 27B workload, Groq 3 LPX reportedly reached 2,529 output tokens per second per user. That kind of headroom gives AI agents more room to perform reasoning steps, execute tool calls, and respond within a practical time budget, even as workloads become more complex.
NVIDIA also shared early feedback around its Vera CPU, which appears to be gaining attention from several AI-focused companies. Early benchmarks from multiple firms showed strong gains in startup speed, latency, throughput, orchestration, and query performance when compared with other CPUs.
Perplexity tested the Vera CPU for its SPACE secure sandbox platform built for agentic AI and reported 1.9x faster sandbox starts. Faster sandbox startup times are important for agentic AI systems because these platforms often need to create isolated environments quickly to run code, test actions, or complete tasks safely.
Daytona also evaluated the Vera CPU for agentic workloads and described the results as significant for agent execution. As AI agents become more capable, the CPU plays a larger role in orchestration, scheduling, data movement, and tool execution, not just traditional compute.
ClickHouse tested Vera CPU performance on ClickBench, its open benchmark for analytical databases. According to the results shared, Vera was the fastest machine the team had measured so far. That is a strong sign for data-intensive workloads, where analytics engines need high memory bandwidth, fast query execution, and consistent latency.
DeepInfra ran several benchmarks on the Vera CPU and reported wins across all measured areas, including 2.2x faster orchestration step latency. Lower orchestration latency can be critical for AI inference pipelines, especially when multiple models, tools, and services must work together in real time.
Prime Intellect found that the Vera CPU maintained high bandwidth and low, consistent memory latency even as more workloads ran in parallel. Predictable performance is essential for agentic AI because fluctuating latency can disrupt multi-step reasoning tasks and reduce overall system responsiveness.
Redpanda reported that its benchmark showed 5.5x lower latencies and 73% higher throughput than other CPUs. Starburst also shared results showing 3x faster query throughput, while Kinetica reported 2.7x faster analytical query performance compared with traditional CPUs.
Taken together, these results suggest that NVIDIA’s Vera CPU is not only aimed at supporting GPUs but also at accelerating the broader infrastructure layer around AI workloads. As AI systems become more complex, CPUs are increasingly responsible for feeding GPUs, managing data flows, coordinating distributed workloads, and supporting fast storage and networking paths.
The larger message from NVIDIA is that the future of AI infrastructure depends on co-design. Silicon, software, networking, storage, security, and power management all need to work together if data centers are going to meet growing AI demand.
That is why NVIDIA is positioning its AI factory architecture as more than a collection of chips. Vera Rubin systems, NVLink-based scale-up computing, Spectrum-X networking, Dynamo inference software, NeMo libraries, BlueField DPUs, and grid-aware power orchestration are all part of the same strategy: extract more useful AI output from every watt.
This approach could become increasingly important as AI workloads shift toward long-context models, real-time inference, multimodal search, autonomous agents, enterprise copilots, and large-scale recommendation systems. These applications require not only raw performance but also efficiency, reliability, and predictable latency.
By focusing on tokens per megawatt, NVIDIA is emphasizing the economics of AI at scale. Higher token throughput within the same power budget can reduce cost per token, improve data center utilization, and allow cloud providers to serve more customers from the same infrastructure.
For enterprises and AI startups, the impact could be significant. More efficient infrastructure may lead to faster response times, lower inference costs, and better access to advanced AI models. For data center operators, it could mean higher revenue potential without waiting years for new grid capacity.
NVIDIA’s Vera Rubin platform and related software stack point toward a future where AI factories are judged not just by how powerful they are, but by how efficiently they turn electricity into useful intelligence. As power becomes one of the defining constraints of the AI boom, every percentage gain in throughput per watt will matter.
The key takeaway is clear: NVIDIA wants to make AI infrastructure smarter from the chip level all the way to the power grid. If these efficiency gains continue to hold across real-world deployments, Vera Rubin and the surrounding AI factory ecosystem could play a major role in shaping the next phase of high-performance, power-aware AI computing.






