NVIDIA Blackwell GPUs Dominate MLPerf Training 6.0 With Record AI Performance
The MLPerf Training 6.0 results are now available, and NVIDIA has once again taken the spotlight in AI training performance. Powered by its Blackwell GPU architecture, the company delivered the fastest results across every benchmark in the latest test suite, reinforcing its lead in large-scale artificial intelligence workloads.
MLPerf Training is one of the most respected benchmark suites for AI hardware because it is open, peer-reviewed, and designed to measure real-world training performance across a range of machine learning models. The newest version adds two important mixture-of-experts tests for modern AI deployments: DeepSeek V3 671B and GPT-OSS 20B.
These additions make the benchmark even more relevant as the AI industry shifts toward larger, more complex models that require massive compute power, high-speed networking, and optimized software stacks.
NVIDIA Blackwell Sets the Pace Across All MLPerf 6.0 Tests
NVIDIA’s Blackwell-based systems, especially the GB300 NVL72 platform, delivered standout results in MLPerf Training 6.0. The company achieved the fastest time to train on every benchmark and was also the only participant to submit results across all seven tests in the suite.
That broad coverage matters. It shows that NVIDIA’s AI platform is not only fast in select workloads but also consistent across different model sizes, architectures, and deployment scenarios.
In the latest MLPerf Training 6.0 results, NVIDIA posted the following times:
DeepSeek V3 671B: 2.02 minutes, with no competing submission
GPT-OSS 20B: 7.43 minutes, with no competing submission
Llama 3.1 405B: 7.07 minutes, with no competing submission
Llama 2 70B LoRA: 0.40 minutes, compared to 8.27 minutes for the nearest alternative
Llama 3.1 8B: 4.46 minutes, compared to 58.63 minutes for the nearest alternative
FLUX.1: 17.1 minutes, compared to 74.44 minutes for the nearest alternative
DLRM-dcnv2: 0.67 minutes, with no competing submission
One of the most striking comparisons came from the Llama 3.1 8B benchmark. NVIDIA completed the workload in 4.46 minutes, while the closest competing system required 58.63 minutes. That represents a performance gap of more than 13 times in time-to-train.
On the newest benchmark additions, including DeepSeek V3 671B and GPT-OSS 20B, rival platforms did not submit results, leaving NVIDIA as the only measured platform for those workloads.
GB300 NVL72 Shows Major Gains Over GB200
NVIDIA’s Blackwell architecture has continued to improve since launch through hardware scaling and software optimization. The GB200 platform already delivered strong performance, but the newer GB300 NVL72 systems push further ahead.
According to the results, GB300 systems can deliver up to 60% higher performance than GB200 in the same NVL72 configuration. A major reason for this improvement is increased AI compute density and support for NVFP4, which helps accelerate AI training workloads while improving efficiency.
This type of progress is important for data centers and cloud providers because training large AI models is expensive, power-hungry, and time-sensitive. Faster training means organizations can iterate more quickly, reduce operating costs, and deploy improved AI models sooner.
Blackwell Scales to 8,192 GPUs in Large AI Training Run
Beyond single-system performance, NVIDIA also demonstrated massive scaling. In MLPerf Training 6.0, Blackwell systems were used in an 8,192-GPU cluster for Llama 3.1 405B training.
The cluster reached the benchmark’s quality target in 7.07 minutes, making it the fastest time-to-train recorded for that test. This result highlights one of NVIDIA’s biggest strengths: combining GPU performance with high-speed interconnects, optimized software, and large-scale system design.
Microsoft Azure scaled Llama 3.1 405B training across 8,192 GPUs using GB200 NVL72 systems and reached the target in 7.07 minutes.
CoreWeave also delivered a major result with DeepSeek V3 671B, reaching the quality target in 2.02 minutes at 8,192-GPU scale using GB300 NVL72 systems connected through Spectrum-X Ethernet networking.
These large-scale results are especially important as AI companies increasingly train models with hundreds of billions of parameters. At that scale, raw GPU power is only one part of the equation. Networking, memory bandwidth, software efficiency, and system reliability all play a major role.
NVIDIA Maintains a Strong Lead Over AMD MI300-Series Accelerators
The MLPerf Training 6.0 results also show NVIDIA’s strong position against AMD’s latest AI accelerators, including MI300-series offerings and newer MI355X-class hardware.
In DeepSeek V3 671B, NVIDIA was the only platform with a submitted result. In FLUX.1, a configuration using 32 NVIDIA GB300 GPUs outperformed much larger accelerator counts from competing AMD-based systems, including configurations with 512 MI300X accelerators and 64 MI325X accelerators.
In Llama 2 70B, NVIDIA’s GB300 and GB200 systems using 8 accelerators also outpaced competing submissions. In Llama 3.1 8B, NVIDIA continued to deliver stronger performance at similar accelerator counts and extended its lead further with larger scale-up configurations.
The results suggest that NVIDIA’s advantage is not limited to having powerful GPUs. Its software ecosystem, networking stack, optimized libraries, and mature AI platform continue to play a central role in real-world training performance.
What MLPerf Training 6.0 Means for the AI Hardware Market
The latest benchmark round makes one thing clear: NVIDIA remains the dominant force in AI training performance. Blackwell GPUs delivered record-setting results across the full MLPerf Training 6.0 suite, while competitors were absent from several of the newest and most demanding tests.
For cloud providers, AI labs, and enterprises building next-generation models, these results reinforce the importance of end-to-end platform performance. Training modern AI models is no longer just about individual chip specifications. It requires tightly integrated systems that can scale from small workloads to thousands of accelerators.
NVIDIA is also preparing its next-generation Vera Rubin AI platform, which is expected to raise performance further in upcoming deployments. But even before that platform arrives, Blackwell is already setting a high bar for the industry.
With record benchmark results, strong scaling to 8,192 GPUs, and continued software-driven performance gains, NVIDIA’s position in AI training remains exceptionally strong.






