NVIDIA has released a detailed technical exploration of its next-generation "Rubin" GPU, including a fully annotated die shot. "Rubin" represents a significant generational leap in NVIDIA's mission to scale computing across numerous racks, systems, and data centers. The company is promoting this generation as the first "agentic AI" native design. This means that inference workloads are continuously active, engaging in reasoning, planning, tool utilization, and executing many sequential steps. In this context, inference refers to a completed model running on the GPUs, rather than being trained, which is what Rubin is designed for. NVIDIA claims up to 10 times more agentic throughput per unit of energy compared to "Blackwell." This improvement relies not only on the new "Vera" CPU and rack-level power management but also on the entire system, as customers now prioritize overall system performance rather than just the performance of a single GPU.
In its full configuration, "Rubin" combines two compute dies in one package through a high-speed link that NVIDIA calls NV-HBI. This chiplet approach uses CoWoS-L packaging from TSMC, chosen because each die is already as large as a chip factory can print in one go, reaching the reticle limit. The GPU contains 336 billion transistors, up to 224 SMs, 896 Tensor Cores with a third-generation Transformer Engine, and 288 GB of HBM4 memory, capable of delivering up to 50 PetaFLOPS of inference and about 35 PetaFLOPS of training at the ultra-low-precision NVFP4 format. This is NVIDIA's proprietary 4-bit data format that enhances efficiency in inference and training with minimal to no accuracy loss compared to the 8-bit and 16-bit data formats used today for training and inference. However, this is quite different from the FP16/FP32 formats used in gaming GPUs. Below, we will break down not just the GPU, but the entire platform NVIDIA has built around it.