NVIDIA yesterday had a GTC conference in Washington D.C., showing its latest "Vera Rubin" Superchip symbiote. Pictured for the first time is the combination of two "Rubin" GPU paired with a single "Vera" CPU carrying 88 custom NVIDIA cores and 176 threads in a single package. NVIDIA quoted performance targets of roughly 50 PetaFLOPS of FP4 compute per Rubin GPU, which yields about 100 PetaFLOPS FP4 for the two-GPU Superchip. The company said engineering samples are already moving through labs and set mass production goals for 2026 with broader shipments and deployments into 2027.
Each Rubin GPU appears to integrate two reticle-sized compute chiplets (2x830mm²?) paired with eight HBM4 stacks, delivering about 288 GB of HBM4 per GPU and roughly 576 GB of HBM4 on the full Superchip. NVIDIA also populated the board with SOCAMM2 LPDDR5X modules to provide large, low-latency system memory, with some older briefings indicating around 1.5 TB of LPDDR5X per Vera CPU on typical trays. The Vera CPU itself uses an 88-core, 176-thread Arm-based custom design and shows signs of a multi-chiplet layout with a distinct I/O chiplet nearby. With "Grace," NVIDIA relied on Arm's Neoverse design, but with Vera, the design team brought this CPU core in the house to extract maximum performance. Additionally, NVLink bandwidth climbs to approximately 1.8 TB/s to sustain heavy CPU-to-GPU traffic in system-demanding workloads such as AI inference and training.