NVIDIA's long-anticipated "Blackwell Ultra" is the final form of the Blackwell family before the transition to "Rubin." However, this silicon doesn't seem to share much resemblance with its predecessors, especially in terms of I/O and performance. On its blog, NVIDIA detailed the silicon design of its latest creation and all the accompanying features like enhanced software support and optimizations. One of the most striking aspects of the Blackwell Ultra, designed for AI servers, is its use of PCIe Gen 6, whereas the consumer Blackwell and regular server Blackwell use PCIe Gen 5. Made using the TSMC 4NP node, the massive Ultra chip features 208 billion transistors, which is 2.6 times more than the last generation Hopper, based solely on raw transistor count. This comes with a 1,400 W TDP, which means a massive cooling system is a must.
When it comes to performance, Blackwell Ultra delivers approximately 1.5 times denser NVFP4 compute compared to Blackwell, resulting in higher tokens-per-second on inference and improved throughput for large-batch training. The chip pairs 160 SMs across two reticle dies via NVIDIA's NV-HBI link, bringing a 10 TB/s die-to-die fabric, 288 GB of HBM3E at up to 8 TB/s bandwidth, and fifth-generation Tensor Cores tuned for NVFP4. Attention-layer throughput benefits from the doubled performance of special function units (SFUs) for transcendental operations, reducing softmax latency and enhancing reasoning responsiveness.