Zen: A New Dawn in High-Performance Computing

At Hot Chips, AMD is delving deep into its latest Zen 5 core architecture, which is set to drive the company’s next phase of high-performance PC development. Initially launching the Zen 1 core architecture in 2017, AMD has since rolled out several advancements, culminating in the recent introduction of Zen 5. This newest architecture brings various enhancements, including a 16% IPC uplift, advanced AVX-512 and FP-512 capabilities, an 8-wide dispatch, 6 ALUs, dual pipe fetch/decode, and utilization of 4nm and 3nm process technologies.

**Key Improvements and Design Objectives:**

AMD outlined the primary design goals for Zen 5 aimed at significantly enhancing 1T and NT performance. The objectives focus on balanced cross-core instruction and data throughput, front-end parallelism, and efficient data movement. Additionally, Zen 5 introduces new ISA extensions, security features, and expanded platform support.

**Core Features Overview:**

– **Threads and Branch Prediction:**
– 2 threads per core.
– Advanced branch prediction with fewer bubbles and increased accuracy.

– **Cache Specifications:**
– I-Cache: 32KB, 8-way; D-Cache: 48KB, 12-way; L2-Cache: 1MB, 16-way.

– **Instruction Fetch and Decode:**
– Dual I-Fetch/decode pipes.
– 4 inst/pipe throughput per pipe in SMT mode.

– **Execution Capabilities:**
– 6 integer ALUs, 4 AGUs, 4 FP ops per cycle.
– Full 512b AVX512 datapaths with new FADD and optimized load/store pipes.

**Execution and Data Flow:**

Zen 5 introduces significant execution and data flow improvements, ensuring:
– Integer execution through 6 ALUs and 4 AGUs.
– Full 512b FP datapaths with 4 execution pipes.
– Enhanced data flow capabilities with 4 load pipes supporting 2, 512b AVX512 pipes, creating higher bandwidth and reduced latency.

**Fetch and Decode Advances:**

– **Branch Prediction Improvements:**
– Enhanced infrastructure with zero-bubble conditional branches, a larger TAGE, and an expanded return address stack.

– **Opcode Cache Efficiency:**
– Improved density and throughput with a 16-way associativity.

**Loading and Storing Enhancements:**

Zen 5’s load and store system has also seen considerable advancements, including:
– 48KB 12-way L1D with 4-cycle load-to-use.
– Higher in-flight window and scalable load ordering queue.
– Optimized data prefetching mechanisms.

**Cache and Configurations:**

Zen 5 has significantly upgraded its cache, doubling the L2/core interface bandwidth and improving L2 and L3 cache performance. Configuration options include:
– Zen 5: Optimized for peak 1T performance.
– Zen 5C: Focused on performance per watt and area efficiency.

**Power Efficiency and ISA Enhancements:**

AMD’s Zen 5 architecture is designed for greater power efficiency, incorporating better branch prediction and optimized operations to minimize bus, cache, and inter-core traffic.

**Product Integration:**

Zen 5 core complexes will be integrated into several upcoming product lines, such as:
– Ryzen 9000 “Granite Ridge” Desktop CPUs.
– Ryzen AI 300 “Strix” Laptop CPUs.
– 5th Gen EPYC “Turin” Data Center CPUs.

AMD’s journey with Zen 5 has just begun, and more products and refinements are anticipated as the architecture matures for both PCs and servers.