NVIDIA Rosa CPU: New Rigel Core Architecture Targets Faster Agentic AI Performance
NVIDIA has shared the first meaningful details about Rosa, its next-generation data center CPU designed to pair with the upcoming Feynman GPU family. Built for the rapidly growing demands of agentic AI, Rosa is expected to push NVIDIA’s CPU roadmap beyond Grace and Vera with a new custom Arm-based core architecture called Rigel.
The Rosa CPU continues NVIDIA’s strategy of building tightly optimized processors for AI systems rather than relying only on GPUs. As AI workloads become more complex, especially in agentic AI environments where models plan, reason, call tools, retrieve data, and execute multi-step tasks, CPU performance becomes increasingly important. Rosa is being positioned as the next step in that evolution.
At the heart of Rosa is Rigel, NVIDIA’s new Arm v9.2 CPU core. Rigel follows the Olympus core used in Vera, but NVIDIA says it will deliver higher per-core performance while keeping the same silicon footprint. That is a key point because it suggests Rosa is not simply about making the chip larger or more power-hungry. Instead, the company is focusing on smarter architecture, better execution efficiency, and stronger single-threaded performance.
Single-threaded CPU performance matters more than ever in AI servers. While GPUs handle massive parallel workloads, CPUs still coordinate many parts of the system. In agentic AI workloads, where applications can involve long chains of decisions and frequent interactions between models, memory, storage, networking, and software layers, faster CPU cores can reduce bottlenecks and improve responsiveness at scale.
NVIDIA says Rigel brings several major improvements over Olympus. These include better instruction delivery, a larger L2 cache, and more efficient memory handling. Each of these upgrades can make a meaningful difference in real-world AI infrastructure.
Better instruction delivery helps keep the CPU cores fed with work, reducing stalls and improving execution efficiency. A larger L2 cache can reduce the need to reach out to slower memory, which is especially important for workloads with frequent data access patterns. More efficient memory handling can also improve performance per watt, a critical factor for hyperscale data centers where power and cooling are major concerns.
Rosa builds on the foundation laid by Grace and Vera. Grace was NVIDIA’s first major modern data center CPU platform, using Arm Neoverse V2 cores and offering strong efficiency for accelerated computing. Vera moved further by introducing NVIDIA’s own Olympus core, based on Arm v9.2, and increased the focus on single-threaded performance. Rosa now takes that philosophy further with Rigel.
Grace features 72 CPU cores, while Vera increases that number to 88 Olympus cores and supports 176 threads through spatial multithreading. NVIDIA has not yet confirmed the core count for Rosa, so it remains unclear whether the company will increase the number of cores or focus primarily on per-core gains. Based on the early details, Rosa appears to prioritize maximum single-thread performance and efficiency rather than simply adding more cores.
The comparison between Grace, Vera, and Rosa shows how quickly NVIDIA’s CPU ambitions are advancing. Grace was designed as an efficient high-core-count processor for accelerated workloads and high-performance computing. Vera was created for AI systems that need high sustained per-core performance and strong memory bandwidth. Rosa is being developed for the next stage: agentic AI systems that require even faster individual cores, lower latency, and better coordination across huge compute clusters.
Memory performance is another major part of the story. Grace uses LPDDR5X memory with ECC and delivers roughly 480 to 512 GB/s of bandwidth per CPU. Vera steps up significantly, offering up to 1.2 TB/s of memory bandwidth and up to 1.5 TB of memory capacity. Rosa’s exact memory configuration has not been disclosed, but it is expected to continue improving efficiency and bandwidth, potentially with future LPDDR6-class memory technologies depending on the final platform.
NVIDIA has also placed heavy emphasis on interconnect performance across its CPU and GPU systems. Grace uses NVLink-C2C for high-speed CPU-to-GPU communication, while Vera improves system connectivity further with a second-generation scalable coherency fabric and faster links. Rosa is expected to continue this trend, although NVIDIA has not yet provided full details on its interconnect design.
The company’s broader roadmap shows Rosa arriving with the Feynman generation, following Blackwell, Rubin, and Vera-based systems. Blackwell is already central to NVIDIA’s current AI platform strategy, while Rubin and Vera are planned to push performance further in the 2026 to 2027 window. Feynman and Rosa are expected to represent the next major leap near the end of the decade.
Based on the current roadmap, Feynman-generation systems are targeted for 2028, with Rosa expected to become part of future data center platforms around that time frame. Wider deployment could follow afterward, with PC-focused variants under future Spark platforms potentially arriving later.
One of the most interesting parts of NVIDIA’s CPU strategy is that the company is not limiting these technologies to the largest data centers. The same architectural direction is expected to influence future workstation and PC-class AI platforms. NVIDIA’s Spark chips are planned to bring combinations of its CPU and GPU technologies into more compact systems, expanding the reach of its AI hardware beyond massive server racks.
The first Spark products are expected to combine Grace and Blackwell technologies, while later generations may move toward Vera Rubin combinations. Rosa and Feynman-based Spark solutions are expected further out, potentially around 2030, if the roadmap remains on track.
Rosa’s importance comes from the changing nature of AI computing. Traditional AI training and inference have been heavily GPU-driven, but agentic AI adds new pressure on CPUs. These workloads often require orchestration, fast decision loops, memory management, data movement, and system-level coordination. A faster CPU core can improve the entire AI pipeline, especially when deployed across thousands or even millions of cores in large data centers.
By developing custom Arm cores such as Olympus and Rigel, NVIDIA is also reducing its dependence on off-the-shelf CPU designs. This gives the company more control over how its processors work with GPUs, memory, networking, and software. That level of vertical integration is becoming increasingly important as AI infrastructure becomes more specialized.
Rosa is still years away from broad availability, and many important specifications remain unknown. NVIDIA has not confirmed the final core count, clock speeds, cache sizes, memory bandwidth, capacity, power targets, or full platform configurations. However, the early details already make one thing clear: Rosa is being designed to extend NVIDIA’s leadership in AI-focused CPU performance.
With Rigel, a larger cache structure, improved instruction delivery, and more efficient memory handling, the NVIDIA Rosa CPU could become a crucial component in the next wave of AI data centers. As Feynman GPUs arrive and agentic AI workloads become more demanding, Rosa may play a key role in keeping future AI systems fast, efficient, and scalable.






