NVIDIA's Vera CPU posts 1.8x per-core gains on agentic workloads, with Perplexity testing production deployment
NVIDIA published benchmark results on July 7 showing its new Vera CPU, built for agentic AI workloads, delivering large per-core performance gains over x86 in real-world testing by Perplexity and other infrastructure partners.
What's new
At the center of Vera is Olympus, NVIDIA's custom CPU core, which the company says "delivers 50% higher instructions per cycle than NVIDIA Grace." The chip pairs that core with up to 1.2TB/s of LPDDR5X memory bandwidth and 3.4TB/s of core-to-core bandwidth across 88 cores, a configuration NVIDIA is pitching specifically at the single-threaded, tool-calling workloads that dominate agentic AI — things like spinning up code sandboxes, running test suites, and executing many small sequential operations rather than the large parallel matrix math GPUs handle.
NVIDIA's headline partner result comes from Perplexity: "Perplexity tested Vera on the agentic work it runs every day." On that workload, NVIDIA reports Vera "delivers 1.8x the sustained per-core performance of x86," and that in Perplexity's own coding-agent tests — cloning repositories and running test suites — Vera "completed the job about 1.5x faster than x86, and started concurrent sandboxes up to 1.9x faster." NVIDIA also cites additional partner numbers: roughly 3x faster large-scale SQL analytics with Starburst, and up to 6x lower latency on real-time streaming with Redpanda.
Context
Vera is NVIDIA's answer to a shift in what AI infrastructure needs to be good at. Grace, NVIDIA's prior CPU line, was built primarily as a high-bandwidth companion to its GPUs for training and batch inference. As agentic systems increasingly spend their time on sequential tool calls — file operations, sandboxed code execution, database queries — rather than pure model inference, single-threaded CPU performance has become a bottleneck that GPU throughput alone doesn't fix. NVIDIA is positioning Vera as the piece of its stack purpose-built to close that gap, rather than a general-purpose server CPU refresh.
Why it matters
Perplexity is a live production AI company, not a benchmarking partner running synthetic tests, so its early results carry more weight than a vendor-run demo — though production deployment is still described as prospective rather than shipped. If the per-core gains hold up at scale, it reinforces a broader trend: as agentic workloads spread through the industry, infrastructure buyers are increasingly optimizing for tool-execution latency and single-thread throughput alongside raw GPU capacity, a shift that also benefits NVIDIA by extending its stack beyond GPUs into a category — CPUs for agent orchestration — it hasn't traditionally competed in against Intel and AMD.
Corroborating sources
- Blogs.nvidia
https://blogs.nvidia.com/blog/nvidia-vera-max-single-threaded-cpu-at-scale/
“Perplexity tested Vera on the agentic work it runs every day.”