Intel Unveils Gaudi3: a challenger to NVIDIA’s Hopper and Blackwell
The company unveils its next-gen AI accelerator.

At Vision 2024, Intel lifted the curtain on the Gaudi3, its next-gen AI accelerator combining two 5nm TSMC dies. The chip boasts impressive specs, featuring 64 Tensor cores (5th Generation), 128GB of HBM2w memory, and 900W of power on air. The Gaudi3 succeeds Gaudi2, developed by Habana Labs, which Intel acquired five years ago.
The Tensor core count has increased from 24 on Gaudi2 to now combining two 32 Tensor cores across two chips. Each chip comes with 48MB of SRAM memory, totaling 96MB per full package. The SRAM memory itself has a bandwidth of 12.8 TB/s. This memory is supported by HBM memory, still relying on HBM2e technology, but with a faster total bandwidth of 3.7 TB/s compared to 2.45 TB/s on Gaudi2. Furthermore, the full capacity has increased from 96GB to 128GB.
Intel Gaudi3 (photograph), Source: Intel
For a PCIe form factor design (HL-388), Gaudi3 is also utilizing the PCI Express 5.0 interface across the full 16 lanes. This version is said to have a TDP of 450 to 600W, a rarely seen figure for this form factor. However, there is also an OAM version (HL-328/325L/335). This version has a TDP rated at 450 to 900W for server air cooling and 900W for the water-cooled edition (HL-335).
Intel Gaudi3, Source: Intel
Intel is making several comparisons to the current AI leader, NVIDIA, who has H100 and H200 chips based on the Hopper architecture and has already announced its successor, Blackwell, which is not expected to launch anytime soon. The current projections indicate that Gaudi3 will be up to 1.7 times faster for training compared to H100 (1.4 to 1.7 times faster depending on the Large Language Model used). However, when it comes to inference speed, the actual performance will vary, sometimes being lower (by up to 10%) and sometimes up to 70% higher. In terms of power efficiency, Gaudi3 is reported to be 1.2 to 2.3 times more efficient, measured with tokens per second per card per watt.
Intel Gaudi3, Source: Intel
Intel is clearly striving to match or surpass NVIDIA, but no figures are provided for AMD Instinct offerings, which have MI300A/X accelerators already announced.
The first samples of Gaudi3 will be provided to partners in the first half of 2024. However, larger volumes should not be expected until the second half of this year. Currently, there are no third-party benchmarks available, so one must rely on Intel’s word for everything.
Intel Gaudi3, Source: Intel
| Intel Gaudi 3 General Specifications | |||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| Family | Part Number | Export Market | Availability | Form Factor | Cooling | TDP [W] | HBM Capacity | Peak HBM Bandwidth | HBM Interface and Type | Last-Level Cache Capacity | Host Interface |
| Gaudi 3-OAM | HL-325L | Non-PRC | Mar ’24 | OAM | Air | 900 | 128GB | 3.7TB/sec | 1024-bit x 8 stacks HBM2e | 96MB | PCIe Gen5 x16 |
| HL-328 | PRC | Jun ’24 | OAM | Air | 450 | 128GB | 3.7TB/sec | 1024-bit x 8 stacks HBM2e | 96MB | PCIe Gen5 x16 | |
| HL-335 | Non-PRC | Oct 24 | OAM | Liquid (1P or 2P) | 900 | 128GB | 3.7TB/sec | 1024-bit x 8 stacks HBM2e | 96MB | PCIe Gen5 x16 | |
| Gaudi 3-PCIe | HL-338 | Non-PRC | Sep ’24 | Dual-slot PCIe, Full Height, 10.5″ Length | Air | 600 | 128GB | 3.7TB/sec | 1024-bit x 8 stacks HBM2e | 96MB | PCIe Gen5 x16 |
| HL-388 | PRC | Sep 24 | Dual-slot PCIe, Full Height, 10.5″ Length | Air | 450 | 128GB | 3.7TB/sec | 1024-bit x 8 stacks HBM2e | 96MB | PCIe Gen5 x16 |
Source: HardwareLuxx , ComputerBase