NVIDIA Announces Rubin AI Platform: 6-Chip Stack, HBM4, and Up to 5x Inference Gains vs Blackwell

NVIDIA used CES 2026 to formally announce its next data center AI platform, Rubin. The stack is built around six new chips and will ship in systems such as the Vera Rubin NVL72 rack and HGX Rubin NVL8 server platform.
At the center is the Rubin GPU, listed at 336 billion transistors and built from two reticle dies. NVIDIA quotes up to 50 PFLOPs of NVFP4 inference and 35 PFLOPs of NVFP4 training, with slides framing that as 5x and 3.5x higher than Blackwell. The GPU moves to HBM4, with up to 288GB per GPU and up to 22 TB/s of memory bandwidth.
Source: NVIDIA
On the CPU side, NVIDIA pairs Rubin with the Vera CPU, listed at 227 billion transistors and based on custom Arm “Olympus” cores. Vera is specified at 88 cores and 176 threads using NVIDIA Spatial Multi-Threading, with up to 1.5TB of LPDDR5x (SOCAMM) and up to 1.2 TB/s of memory bandwidth. Coherent NVLink-C2C bandwidth is listed at 1.8 TB/s, and NVIDIA claims 2x gains in data processing, compression, and CI/CD versus Grace.
Source: NVIDIA
For scale up, NVLink 6 is rated at 3.6 TB/s of bidirectional GPU to GPU bandwidth per GPU, and NVIDIA quotes 260 TB/s of scale up bandwidth for the NVL72 rack. The NVLink 6 switch is also positioned as part of the compute fabric, with figures citing 28.8 TB/s total bandwidth and 14.4 TFLOPS of FP8 in network compute per switch tray. For scale out, NVIDIA is rolling ConnectX-9 and BlueField-4, with ConnectX-9 called out at up to 1.6 Tb/s per Rubin GPU and BlueField-4 positioned as an 800 Gb/s DPU, while Spectrum-X Ethernet Photonics is tied to Spectrum-6 and a 102.4 Tb/s switch infrastructure with co-packaged optics.
Source: NVIDIA
At the rack level, NVIDIA’s flagship configuration is the Vera Rubin NVL72, described as 72 Rubin GPUs and 36 Vera CPUs connected through NVLink 6. NVIDIA’s own figures cite 3.6 EFLOPS of NVFP4 inference and 2.5 EFLOPS of training, plus 20.7TB of HBM4 capacity and 54TB of LPDDR5x capacity, along with 1.6 PB/s of HBM bandwidth. On efficiency, NVIDIA is tying Rubin to large drops in AI cost metrics, including up to 10x lower inference token cost and 4x fewer GPUs for MoE training versus Blackwell, with one slide also framing MoE inference at about one seventh the token cost versus GB200.
NVIDIA says Rubin is already in full production in Q1 2026, despite earlier guidance pointing to mass production in the second half of 2026. Partner availability is still described as the second half of 2026, and NVIDIA lists early 2026 deployments across AWS, Google Cloud, Microsoft, and Oracle Cloud, along with NVIDIA Cloud Partners such as CoreWeave, Lambda, Nebius, and Nscale.
| NVIDIA Vera vs Grace CPUs | ||
|---|---|---|
| VideoCardz.com | Grace CPU | Vera CPU |
| Cores | 72 Neoverse V2 cores | 88 NVIDIA Custom Olympus cores |
| Threads | 72 | 176 Spatial Multi-Threading |
| L2 Cache per core | 1MB | 2MB |
| Unified L3 Cache | 114MB | 162MB |
| Memory bandwidth (BW) | Up to 512GB/s | Up to 1.2TB/s |
| Memory capacity | Up to 480GB LPDDR5X | Up to 1.5TB LPDDR5X |
| SIMD | 4x 128b SVE2 | 6x 128b SVE2 FP8 |
| NVLINK-C2C | 900GB/s | 1.8TB/s |
| PCIe/CXL | Gen5 | Gen6/CXL 3.1 |
| Confidential compute | NA | Supported |
Source: NVIDIA