AMD MI455X uses twelve HBM4 stacks for 432GB memory capacity
© AMD AMD has introduced the Instinct MI455X, its CDNA 5 data center GPU for large AI training, inference and rack-scale systems. The accelerator includes 432GB of HBM4 memory and offers up to 23.3 TB/s of peak memory bandwidth.
AMD lists 320 billion transistors and a 2nm process node. The package combines several process technologies rather than using one node for every chiplet.
Eight compute dies and twelve HBM4 stacks
The MI455X uses eight Accelerator Complex Dies manufactured on TSMC N2. These chiplets provide 256 active Work Group Processors. AMD also uses two N3P Fabric and Cache Dies and two N3P I/O Dies.
Twelve HBM4 stacks connect through a 192-channel memory interface. The GPU has 192MB of global L2 cache, split into two 96MB sections. Its I/O configuration supports two PCIe Gen 6 links or three AMD AI-NICs using UALink.
Up to 40.26 PFLOPS in MXFP4
AMD rates the MI455X for 20.13 PFLOPS of peak MXFP8 and MXFP6 compute. MXFP4 performance reaches 40.26 PFLOPS. Compared with the Instinct MI355X, AMD claims up to 1.5 times more memory, 2.9 times more memory bandwidth and four times higher MXFP8 and MXFP4 throughput.









CDNA 5 adds a Tensor Data Mover for direct asynchronous transfers between global memory and local data storage. AMD has also added work-group clustering, multicast transfers, data prefetching, split barriers and a new command processor intended to reduce dispatch latency.
The MI455X can operate as one, two or up to eight spatial partitions. NPS1 interleaves addresses across all twelve HBM4 stacks, while NPS2 divides the GPU into two domains with six stacks each and avoids transfers between the two main die groups.
AMD also announced the Instinct MI430X for scientific computing, sovereign AI and combined AI-HPC workloads. AMD lists up to 288 TFLOPS of hardware FP64 performance for that model, but has not disclosed its full memory configuration or pricing.
Source: AMD