NVIDIA Vera CPU Gains Momentum as AI Firms Look Beyond GPUs for Agentic Workloads
NVIDIA’s Vera CPU is quickly becoming one of the most closely watched chips in the AI infrastructure market. While GPUs remain the backbone of modern AI training and inference, a new wave of agentic AI workloads is increasing demand for CPUs that can deliver stronger single-threaded performance, predictable latency, and faster response times.
The Vera CPU appears to be designed exactly for that shift. NVIDIA is positioning the chip as an inference-optimized processor built for the next generation of AI systems, where speed, responsiveness, and per-core efficiency matter just as much as raw parallel compute.
Perplexity is among the latest AI companies to show interest in NVIDIA’s Vera CPU. Nate Kupp, Vice President at Perplexity, said Vera stood out as a strong match for many of the company’s core workloads. According to reported comments, the chip delivered around 1.5 times faster performance than traditional CPUs in agentic AI coding tasks.
That kind of gain is significant because agentic AI systems work differently from conventional AI workloads. Instead of simply processing one request and returning an answer, AI agents often operate in repeated loops. They interpret a task, execute a step, review the result, decide what to do next, and continue the process until the objective is complete.
In this type of workflow, every delay adds up. The faster each CPU core can complete its part of the loop, the more responsive the AI system becomes. This is where NVIDIA believes Vera has an advantage.
NVIDIA describes Vera as a “Max Single-Threaded” CPU at scale. In simple terms, that means the chip is built to deliver strong performance from each individual core rather than relying only on adding more cores. For agentic AI, this can be more valuable than simply increasing core count.
The company says an ideal CPU for these workloads needs three things: strong per-core performance under load, enough memory bandwidth to keep each active core supplied with data, and predictable latency. Without those qualities, adding more cores may not improve performance in the areas that matter most. In some cases, too many cores competing for shared resources can even create bottlenecks.
Vera takes a different approach. Instead of chasing the highest possible core count, NVIDIA has focused on making each core faster and more efficient. The CPU uses custom Olympus cores and is said to deliver a 50% IPC improvement over Grace, NVIDIA’s previous data center CPU platform.
Memory bandwidth is another key part of Vera’s design. The chip supports up to 1.2 TB/s of LPDDR5X memory bandwidth while keeping memory power under 40 watts. It also uses a monolithic compute die designed to keep active cores fed with data and reduce unpredictable delays.
NVIDIA claims Vera offers 3.4 TB/s of core-to-core bandwidth, which the company says is more than three times faster than other data center CPU offerings. This is intended to help all 88 cores access the full memory performance of the CPU without creating major slowdowns.
For AI companies, this matters because inference workloads are becoming more complex. As AI models become more interactive, autonomous, and capable of multi-step reasoning, the CPU plays a larger role in coordinating tasks, handling logic, managing data movement, and supporting real-time decision-making.
Perplexity reportedly saw a 50% uplift in agentic AI workflows compared with traditional x86 CPUs. In concurrent sandbox environments, performance gains were said to reach up to 90%. NVIDIA has also highlighted results from partners showing up to 3x faster performance in large-scale SQL analytics with Starburst and up to 6x lower latency in real-time streaming with Redpanda compared with x86-based systems.
These numbers suggest that Vera is not just aimed at AI model serving, but also at the broader data infrastructure surrounding AI applications. As companies build more advanced AI products, they need faster databases, lower-latency streaming systems, and CPUs that can keep up with GPU-driven inference pipelines.
NVIDIA’s broader ambition for Vera is also substantial. The company expects the CPU to become a major revenue driver, with projections reaching as high as $20 billion. Vera is already said to be in mass production and has reportedly attracted interest from major AI and cloud companies, including OpenAI, xAI, Oracle, and Anthropic.
The growing attention around Vera reflects a larger industry trend. AI hardware is no longer only about building bigger GPU clusters. As agentic AI, real-time inference, AI coding tools, and autonomous assistants become more important, CPUs optimized for low-latency, single-threaded execution are becoming a critical part of the stack.
Competition in the data center CPU market is also intensifying. Traditional x86 vendors continue to improve inference performance, while major AI companies are exploring custom silicon for their own internal workloads. NVIDIA’s advantage is that it can integrate CPUs, GPUs, networking, and software into a broader AI platform, making Vera part of a larger ecosystem rather than a standalone processor.
The company is already looking beyond Vera as well. NVIDIA has started discussing its next-generation data center CPU, codenamed Rosa, which is expected to use an updated Rigel core architecture. While Vera is designed to establish NVIDIA’s stronger presence in AI-focused CPUs, Rosa could push that strategy even further.
For now, Vera’s appeal comes from a simple idea: AI agents need speed at every step. In workloads where decisions happen in rapid loops, nanoseconds matter. NVIDIA is betting that a CPU with stronger per-core performance, high memory bandwidth, and predictable latency can become essential for the next era of AI computing.
If demand from companies like Perplexity continues to grow, Vera could mark a major turning point in how AI infrastructure is built. GPUs may still dominate AI acceleration, but CPUs optimized for agentic inference could become just as important in delivering fast, responsive, and scalable AI experiences.






