During the AI Infra Summit, NVIDIA announced the "Rubin CPX" GPU, a specialized accelerator derived from the upcoming "Rubin" family, specifically made for massive-context AI models. The chip delivers 30 PetaFLOPS of NVFP4 compute performance on a monolithic die, accompanied by 128 GB of GDDR7 memory. The monolithic die configuration represents a departure from the dual-GPU packages characteristic of NVIDIA's current Blackwell and Blackwell Ultra architectures, as well as the design path that the rest of the Rubin family will follow. The Rubin CPX addresses computational bottlenecks in extended-context scenarios where AI models process millions of tokens simultaneously. This capability proves critical for applications including comprehensive software codebase analysis and hour-long video content processing, which can require up to one million tokens.
The processor integrates four NVENC and four NVDEC video encoders directly on-chip, enabling streamlined multimedia workflows without external processing dependencies. Performance metrics suggest that the Rubin CPX is delivering three times the attention processing speed of NVIDIA's current best GB300 Blackwell Ultra accelerator systems. The architecture employs a cost-optimized single-die approach, rather than multi-chip modules, which potentially reduces manufacturing complexity while maintaining computational density. Memory bandwidth specifications remain undisclosed, though a 512-bit interface could yield approximately 1.8 TB/s throughput when using 30 Gbps GDDR7 memory chips. NVIDIA plans integration of Rubin CPX processors within the Vera Rubin NVL144 CPX platform, which combines traditional Rubin GPUs with the specialized CPX variants. This hybrid configuration targets 8 ExaFLOPS of aggregate compute performance, with 1.7 PB/s of memory bandwidth, across a complete rack deployment. The "Kyber" rack will include ConnectX-9 network adapters capable of 1600G networking, Spectrum6 doing 102.4T switching, as well as co-packaged optics. NVIDIA plans this to arrive in late 2026, after the regular Rubin GPU launch in early 2026.