NVIDIA Vera dual-socket platform offers 176 PCIe Gen 6 lanes

NVIDIA has published new architectural details for its Vera data center CPU and the custom Olympus core used by the processor. Vera combines 88 Olympus cores, 176 hardware threads and a 164MB unified L3 cache on a monolithic compute die.
Each Olympus core uses a wide front end with a neural branch predictor, a 64KB four-way L1 instruction cache and a 48-instruction decode queue. The core can fetch up to 16 instructions per cycle and includes a 10-wide decoder capable of processing ten fused instructions per cycle.
© NVIDIA The mid-core can rename, allocate and commit up to ten micro-operations per cycle. NVIDIA also lists memory renaming, value prediction and move elimination among the technologies used to reduce dependency stalls. The execution engine includes integer, branch, vector, floating-point, cryptography and dedicated load/store resources.









Each core has a 96KB six-way L1 data cache and a 2MB eight-way L2 cache. NVIDIA has also added several hardware prefetch mechanisms, including a graph prefetcher for pointer-heavy data structures and graph-processing workloads.
Vera uses NVIDIA Spatial Multithreading, which provides two hardware threads per Olympus core. NVIDIA says the design can partition core resources between the threads to reduce contention compared with conventional simultaneous multithreading. A core can prioritize one performance-sensitive thread while its second thread handles system and management tasks.
© NVIDIA The CPU connects its cores, cache, memory controllers and I/O through the second-generation NVIDIA Scalable Coherency Fabric. NVIDIA lists up to 3.4 TB/s of core-to-core bandwidth, 1.2 TB/s of SOCAMM2 LPDDR5X bandwidth and up to 1.5TB of memory capacity per CPU. Vera also supports an 1.8 TB/s NVLink-C2C interface, PCIe Gen 6 and CXL 3.1. Dual-socket systems provide 176 PCIe lanes and use a two-node NUMA configuration, with one NUMA domain per socket.
NVIDIA claims Vera delivers up to 1.8 times higher performance than unnamed x86 systems in selected agent workloads. The results are based on NVIDIA’s internal SPEC CPU 2026 testing from July 2026 and have not yet been independently verified.
Source: NVIDIA