Introducing large maximum single-threaded CPUs – a revolutionary category designed for the agent AI era.
In the development and deployment of AI agents, the CPU becomes essential for inference, response time, and learning. It executes tasks mandated by AI models, such as tool invocation, code execution, data processing, key-value caching, and result analysis.
For AI factory agents, speed is critical.
A faster CPU results in quicker task execution by the agent, driving productivity.
In an AI factory, maximizing GPU utilization is vital. Delays in CPU task completion hinder revenue and negatively affect GPU efficiency. Thus, having a CPU with superior single-threaded performance is crucial for optimizing both revenue and agent performance.
Current data center CPUs often lack the speed required.
While fast CPUs exist for PCs and workstations, data center CPUs have shifted focus from single-threaded performance. The rise of cloud computing has led manufacturers to prioritize higher core counts, sacrificing individual performance to reduce costs.
CPUs optimized for cost per core have increased core counts but compromised on features that enhance speed, like high-performance memory fabric and efficient instruction processing per core. The shift to chiplet architecture has further minimized costs but introduced a “chiplet tax,” hindering optimal memory performance.
AI agents demand CPUs designed to maximize single-threaded performance at scale.
Large maximum single-threaded CPUs ensure fast processing of each agent step, even under full system load. Each core exhibits peak performance without interference, tailored to deliver:
- Exceptional per-core performance under load
- Adequate memory bandwidth per core to maintain data flow
- Predictable and low latency
Each core can efficiently complete tasks without delays from neighboring cores, leading to enhanced throughput and optimal single-core performance.
NVIDIA Vera exemplifies this advanced CPU design.
Building a High-Performance Single-Threaded CPU for Agent Loops
AI agents operate in continuous loops rather than completing isolated tasks. They make inferences for subsequent actions, with the CPU executing these tasks and returning results—then iterating through the cycle.
This continuous model differs significantly from traditional CPUs, which are optimized for sporadic, user-driven interactions. Agent tasks are persistent and concurrent, requiring executors that function seamlessly together.
Increased core counts on a CPU translate to more simultaneous agent tasks; however, this doesn’t hasten individual processing within an agent’s loop. Essentially, adding cores doesn’t expedite single-thread performance and can even diminish it due to resource competition.
Individual core performance is critical for expediting task completion within agent loops. While the aggregate processing capability of additional cores is beneficial, it’s not the sole measure of effectiveness. Given that actions depend on prior results, the performance per core significantly influences loop velocity.
Ultimately, the leading CPUs for agents must excel in single-threaded performance, maintaining efficiency across all cores. In agent processing, every nanosecond counts, rendering NVIDIA Vera a standout in this evolving landscape.
NVIDIA Vera: The Premier Single-Threaded CPU for Agents
NVIDIA Vera is a pioneering single-threaded CPU architectured specifically for agent loops, facilitating the necessary operations between model calls, particularly involving tool use, data processing, coding execution, and result evaluation.

At Vera’s core is Olympus, NVIDIA’s custom CPU core that achieves 50% more instructions per cycle than NVIDIA Grace. This increased efficiency is vital for sequential agent tasks such as tool calls and data processing, as prompt completion drives quicker results in subsequent model calls.
Vera’s fast cores are complemented by a staggering 1.2 TB/s of LPDDR5X memory bandwidth at under 40 watts of power consumption. Its monolithic compute die ensures that all active cores remain well-supplied with data, along with 3.4 TB/s inter-core bandwidth—three times that of competing data center CPUs—thus avoiding performance bottlenecks.
This combination results in accelerated agent loops. When subjected to agent workload simulations, Vera offers 1.8 times more sustained per-core performance than x86-based systems.
Enhanced performance compounds through multiple tasks—including tool use, data processing, and validation—optimizing AI factory operations while utilizing existing GPU resources effectively.
Vera’s performance was tested through practical coding tasks (like repository cloning and test execution), achieving approximately 1.5 times the speed of x86, and concurrently starting up to 1.9 times faster. Notably, Perplexity is considering Vera for its upcoming operational platform.
Agents are also data-intensive. Vera processes CPU-side data tasks swiftly, continuously querying, filtering, and managing information. Partners have observed large-scale SQL analytics executing three times faster with Starburst, along with six times lower latency in real-time streaming using Redpanda, in comparison to leading x86 CPUs.
Agent workflows are multifaceted, encompassing tool execution, data processing, and reinforcement learning for model training—all unified by Vera’s strengths.
Rather than relying on various CPUs for diverse tasks, a single Vera can efficiently manage all workloads. Furthermore, it supports NVIDIA Vera Rubin GPUs and powers the NVIDIA BlueField-4 STX storage processor, allowing AI factories to function on a single architecture.
NVIDIA’s innovation isn’t stopping here. The forthcoming Rosa CPUs featuring Rigel cores will extend NVIDIA’s CPU roadmap for the agent AI era, delivering higher performance per core without increasing the silicon footprint. Key enhancements will include improved instruction delivery, increased L2 cache, and superior memory management.
Engineered for Agent Speed
As we advance into the agent AI age, billions of agents will depend on efficient CPUs to operate, check, retrieve, execute, and validate tasks. In this burgeoning market, the output of agents becomes the primary product. Quicker agent loops provide GPUs more time to generate revenue while minimizing latency.
NVIDIA Vera is crafted for this future.
For additional insights, visit NVIDIA Vera CPU.
Source: blogs.nvidia.com


