BlogsNVIDIANVIDIA Vera Storage Processor Performance

NVIDIA Vera Storage Processor Performance

NVIDIA Vera Storage Processor Performance

3
posts
2026

This feature thread tracks the evolution of NVIDIA's storage processing capabilities, focusing on the performance and efficiency gains provided by the NVIDIA Vera BlueField-4 STX Storage Processor. Initial efforts focused on understanding the demands of AI-native storage for agentic AI workloads, which require high throughput for operations like encryption, compression, integrity checking, and data recovery. Subsequent developments have demonstrated the Vera CPU's ability to significantly outperform x86-based architectures in agentic workloads, particularly in sandboxed execution, by providing high per-core performance, efficient memory bandwidth, and a unified coherency fabric. This enables faster task completion, improved accelerator utilization, and increased overall AI factory output.

2026

NVIDIA Vera Storage Benchmarks: Faster Encryption, Compression, Integrity Checking, and Recovery for AI-Native Storage | NVIDIA Technical Blog

8/3/2026

This post introduces and benchmarks the NVIDIA Vera BlueField-4 STX Storage Processor, highlighting its performance advantages over x86 CPUs for AI-native storage workloads. It details the Vera CPU architecture, including its 88 Armv9.2 cores, Spatial Multithreading, Scalable Coherency Fabric, and SOCAMM2 LPDDR5X memory, and quantifies its throughput gains in encryption (up to 1.43x), decryption (up to 1.29x), Reed-Solomon recovery (up to 3.26x), CRC32C integrity checking (up to 3.67x), compression (up to 3.29x), decompression (up to 1.72x), and multi-stage pipeline operations (up to 3.21x). The post emphasizes how these improvements enable higher service density and more concurrent data flows for agentic AI workloads by reducing CPU, power, and cooling overhead in the storage data path.

NVIDIA Vera CPU Boosts AI Factory Throughput to Accelerate Agentic Workloads | NVIDIA Technical Blog

7/7/2026

This post details how the NVIDIA Vera CPU architecture, with its monolithic compute die, unified cache, and Scalable Coherency Fabric, delivers 1.8x faster sustained per-core performance under full socket load, directly improving reinforcement learning (RL) training throughput, policy gradient quality, and environment rollout completion rates. It achieves 40% lower peak loaded latency and over 3x the per-core memory bandwidth at less than half the power of traditional x86 data center CPUs, optimizing agentic inference and interactive deployment responsiveness at scale. By minimizing CPU-side stalls, reducing context reconstruction from KV-cache evictions, and maximizing throughput, NVIDIA Vera CPU enables higher GPU utilization and efficiency in densely loaded agentic AI factories, directly increasing overall system productivity and service level agreement adherence.

NVIDIA Vera CPU Sets a New Standard for Agentic Workloads in AI Factories | NVIDIA Technical Blog

6/1/2026

This post introduces the NVIDIA Vera CPU as a new design point for AI factories, specifically addressing the increased importance of CPU performance in agentic AI and reinforcement learning workloads. It details how the Vera CPU, with its 88 NVIDIA Olympus cores and 1.2 TB/s LPDDR5X memory bandwidth, delivers high per-core performance, concurrency, and energy-efficient memory bandwidth. Key architectural features like the neural branch predictor, wide decode unit, deep out-of-order engine, graph prefetcher, and NVIDIA Scalable Coherency Fabric are highlighted for their role in accelerating agentic tasks such as sandboxed code execution, data retrieval, and orchestration. Performance benchmarks show over 1.8x higher agentic sandbox performance compared to x86-based architectures.