NVIDIA Blackwell GB300 Sets New Benchmark Record for Agentic AI Performance
NVIDIA’s Blackwell Ultra GB300 platform has delivered a major performance milestone in a new AI benchmark designed to measure real-world agentic AI workloads. In the latest AA-AgentPerf results from Artificial Analysis, the GB300 NVL72 system achieved the highest recorded performance so far, showing a massive generational leap over Hopper-based hardware.
AA-AgentPerf is a new benchmark focused on how well AI inference systems handle agentic AI, the type of workload used by modern AI assistants, coding agents, and automated reasoning tools. Unlike traditional benchmarks that rely on simple prompts or isolated requests, AA-AgentPerf evaluates continuous, multi-step AI activity that more closely reflects how advanced AI systems are used in production.
The benchmark measures realistic agent behavior, including multi-turn coding tasks, reasoning steps, tool usage, and changing context lengths. It also tests sustained concurrent workloads, where many active agents are running at the same time. This puts pressure on important parts of an AI deployment, including KV cache reuse, scheduling efficiency, speculative decoding, latency, and overall throughput.
Artificial Analysis also uses market-based service-level targets to reflect performance expectations seen across real-world serverless AI providers. This makes the benchmark especially relevant for companies building or scaling AI infrastructure for demanding agentic workloads.
The benchmark focuses on three major performance areas: time to first token, output speed, and total system output throughput. Time to first token measures how quickly the system begins responding after receiving a request. Output speed measures how many tokens per second are generated after the first token appears. System output throughput measures the total output across all active AI agents running at the same time.
For its first published AA-AgentPerf results, NVIDIA tested the DeepSeek V4 Pro model on the GB300 NVL72 platform. This type of frontier AI model is representative of the large-scale models now powering AI agents, coding assistants, and enterprise AI workflows.
The results show a dramatic advantage for Blackwell Ultra. NVIDIA says the GB300 NVL72 can support up to 61,400 concurrent agents per megawatt, compared to around 2,600 concurrent agents per megawatt on the HGX H200 platform. That gives GB300 roughly a 20x efficiency lead over the previous Hopper generation in this specific agentic AI benchmark.
Hardware efficiency also improved significantly. The GB300 NVL72 reached 57.5 concurrent agents per GPU, while the H200 delivered around 1.4 concurrent agents per GPU. This indicates that Blackwell is not only delivering higher raw performance, but also using its GPUs far more effectively when handling many simultaneous AI agent sessions.
These results highlight one of Blackwell’s key strengths: keeping GPUs highly utilized during complex, concurrent AI workloads. Agentic AI is more demanding than standard text generation because requests are often long, irregular, and multi-step. A single AI agent may need to reason, call tools, write code, inspect outputs, and continue working through several turns. Scaling that behavior to tens of thousands of simultaneous agents requires both powerful hardware and highly optimized software.
The GB300 NVL72 platform appears designed for exactly this kind of workload. By improving throughput, latency, and energy efficiency, NVIDIA is positioning Blackwell Ultra as a major platform for large-scale AI inference, especially as businesses move from simple chatbot interactions to more autonomous AI agents.
Looking ahead, NVIDIA’s upcoming Rubin architecture is expected to push AI performance even further. Rubin is planned to bring a new AI architecture with major compute improvements, including up to 50 PFLOPs of NVFP4 compute. Paired with the Vera CPU, Rubin could improve large language model tool calls, end-to-end AI responsiveness, and overall efficiency in future agentic AI deployments.
The latest AA-AgentPerf results suggest that the AI hardware race is moving beyond raw training performance and into large-scale inference efficiency. As AI agents become more common in coding, research, productivity, customer support, and enterprise automation, the ability to support thousands of active agents efficiently may become one of the most important measures of next-generation AI infrastructure.






