A person in a red jacket holding a chip on stage and another person in a black leather jacket holding a motherboard with a visible logo.

MLPerf 6.1 Showdown: AMD’s 512-GPU MI355X Cluster Challenges NVIDIA Blackwell as Intel Joins the Race

MLPerf Inference v6.1 Highlights AMD, NVIDIA and Intel in the Race for AI Performance

MLCommons has released MLPerf Inference v6.1, giving the AI hardware industry a fresh look at how today’s most powerful accelerators, CPUs and workstation GPUs perform across modern inference workloads. The new benchmark round reflects where the market is heading: larger language models, faster token generation, better server throughput, and stronger scaling across multi-GPU systems.

This latest set of results includes submissions from AMD, NVIDIA and Intel, covering everything from massive data center clusters to workstation-class hardware. The outcome is not a simple one-company victory. Instead, MLPerf v6.1 shows a fast-moving AI hardware landscape where performance depends heavily on model type, system size, software maturity and scaling efficiency.

AMD pushes scale with 512 Instinct MI355X GPUs

One of the biggest highlights of MLPerf Inference v6.1 is AMD’s use of its largest submitted cluster configuration to date: 512 Instinct MI355X GPUs. This system delivered leading performance in both offline and server scenarios for DeepSeek R1, showing AMD’s growing strength in large-scale AI inference deployments.

The same 512-GPU configuration also performed strongly in GPT-OSS 120B, where AMD again led in both offline and server use cases. These results point to AMD’s increasing focus on production-scale AI workloads, where raw throughput and efficient scaling are critical.

At 72 GPUs, AMD reported 95% scale efficiency on GPT-OSS 120B with the Instinct MI355X platform. In a separate large-scale submission, Crusoe achieved 5.75 million offline tokens per second on GPT-OSS 120B and 2.90 million offline tokens per second on DeepSeek R1 using 512 GPUs. That level of aggregate token throughput shows how quickly large AI clusters are evolving for real-world inference.

NVIDIA GB300 and Vera Rubin make a strong impression

NVIDIA’s Grace Blackwell Ultra GB300 platform also made a major showing in MLPerf Inference v6.1. In 288-GPU configurations, GB300 delivered results close to AMD’s top MI355X cluster in DeepSeek R1 and showed strong scaling across large deployments.

One of NVIDIA’s most notable claims is that a 288-GPU GB300 NVL72 submission across four racks reached 99% scaling efficiency. That means throughput increased almost linearly from a single-rack baseline, an important achievement for organizations building large AI infrastructure.

Even more eye-catching was the early MLPerf preview of NVIDIA’s Vera Rubin VR200 systems. In the submitted results, a 72-GPU Vera Rubin NVL72 configuration was up to 95% faster than a 72-GPU GB300 setup. Meanwhile, a 36-GPU Vera Rubin configuration managed to outperform a 72-GPU GB200 configuration.

Because this is still an early preview, the results are especially important. They suggest that Vera Rubin could bring a major generational leap once it becomes more widely deployed and optimized for future MLPerf rounds.

GB300 leads in several smaller and mid-size GPU configurations

While AMD’s largest 512-GPU clusters stood out at extreme scale, NVIDIA’s GB300 showed strong performance at more common 8-GPU and 72-GPU system sizes.

In GPT-OSS 120B, GB300 Grace Blackwell Ultra led the 72-GPU and 8-GPU categories. For Llama 2 70B, NVIDIA’s Blackwell GPUs dominated the 72-GPU configurations. GB300 also outperformed AMD’s MI355X in 8-GPU Llama 2 70B tests, although AMD’s Instinct platform still came in ahead of NVIDIA’s GB200 in some comparisons.

For Llama 3.1 8B, an 8-GPU NVIDIA GB300 configuration was up to 17% faster than an 8-GPU AMD MI355X configuration. That gives NVIDIA a strong position in smaller language model inference, especially in deployments where fewer GPUs are preferred for cost, power or space reasons.

AMD MI350P shows progress over MI300X

AMD also submitted results for the Instinct MI350P, which delivered better performance than the MI300X in 8-GPU configurations. That improvement gives AMD another competitive option in AI inference, especially for organizations looking at newer Instinct hardware but not necessarily deploying massive clusters.

AMD said continued ROCm software optimization helped improve performance on the same MI355X GPU hardware within a single MLPerf cycle. According to the company, these updates increased GPT-OSS 120B throughput on 8-GPU systems, reduced Wan-2.2 latency, and improved cluster throughput with fewer GPUs.

This is a key theme in MLPerf v6.1: hardware matters, but software optimization is just as important. Better compilers, libraries, kernels and runtime improvements can significantly raise performance without changing the silicon.

Intel expands Xeon and Arc Pro participation

Intel also had a meaningful presence in MLPerf Inference v6.1, with both Xeon 6 server CPUs and Arc Pro B70 workstation GPUs included in the benchmark submissions.

On the CPU side, Intel expanded its Xeon 6 participation from two benchmarked SKUs in MLPerf v6.0 to five in v6.1. The number of CPU inference results increased from 24 to 35. Intel also remains the only company with a standalone server CPU represented in the MLPerf Inference submissions.

That is notable because most AI inference discussions focus heavily on GPUs and accelerators. Intel’s Xeon results show there is still a role for CPUs in certain inference workloads, particularly where flexibility, system simplicity or CPU-based deployment is important.

Intel’s Arc Pro B70 also appeared across several AI workloads, including Llama 3.1 8B, Llama 2 70B, GPT-OSS 120B, Whisper and end-to-end retrieval-augmented generation. A four-GPU Arc Pro B70 system provides 128GB of VRAM, making it an interesting workstation-class option for AI developers and smaller deployments.

Intel reported that, on the same four-GPU Arc Pro B70 system used in the previous benchmark round, GPT-OSS 120B server performance improved by 36%, while offline performance rose by 27%. Those gains highlight the impact of Intel’s ongoing software stack improvements.

Workstation GPUs are becoming more relevant for AI inference

Another important takeaway from MLPerf v6.1 is the growing role of workstation-class GPUs. NVIDIA RTX PRO GPUs and Intel Arc Pro B70 appeared in several benchmark categories, including Llama 2 70B, Llama 3.1 8B and GPT-OSS 120B.

This matters because not every AI workload runs in a hyperscale data center. Many businesses, researchers, developers and content creation teams need local AI performance in workstations or smaller servers. The appearance of these platforms in MLPerf suggests the AI hardware market is becoming broader, with more options beyond top-end data center accelerators.

Software optimization continues to reshape the results

The most important lesson from MLPerf Inference v6.1 may be that benchmark numbers are constantly changing. AI hardware performance is no longer defined only at launch. Continuous software improvements can dramatically increase throughput, reduce latency and improve power efficiency over time.

NVIDIA said software optimizations in its MLPerf Inference v6.1 submissions delivered up to 1.6 times higher performance compared with v6.0. The company also said additional optimizations after the v6.1 submission window produced even more gains.

AMD reported improved results from ROCm optimization on the same Instinct MI355X hardware within one MLPerf cycle. Intel also showed notable improvements on Arc Pro B70 without changing the underlying four-GPU system.

This means the MLPerf v6.1 results should be viewed as a snapshot rather than a final ranking. As software stacks mature, drivers improve and AI frameworks become more efficient, the same hardware can deliver better results in future benchmark rounds.

What MLPerf Inference v6.1 tells us about the AI hardware market

MLPerf Inference v6.1 shows an increasingly competitive AI hardware industry. AMD is demonstrating strong large-scale performance with massive Instinct MI355X clusters and improved ROCm software. NVIDIA continues to lead in several 8-GPU and 72-GPU categories with GB300, while its early Vera Rubin preview points to a major upcoming leap. Intel is expanding its presence with Xeon 6 CPUs and Arc Pro B70 GPUs, proving that CPU and workstation AI inference still have a place in the benchmark conversation.

There is no single winner across every model and every configuration. AMD shines at extreme scale and in several GPT-OSS 120B and DeepSeek R1 scenarios. NVIDIA remains highly competitive across Blackwell systems and appears to be preparing a major step forward with Vera Rubin. Intel is building momentum through broader submissions and software-driven gains.

For AI developers, cloud providers, enterprise buyers and hardware enthusiasts, MLPerf v6.1 makes one thing clear: the AI inference race is accelerating. Performance is improving not only through new chips, but also through better software, smarter scaling and more efficient system design.

The next MLPerf round could look very different. As Vera Rubin systems become more widely tested, as AMD continues refining ROCm, as NVIDIA pushes CUDA and platform-level optimizations, and as Intel improves its AI software stack, the benchmark landscape will keep shifting. That rapid progress is exactly what the AI industry needs as demand for faster, more efficient inference continues to grow.