Modern AI workloads demand more than raw compute—they require a tightly integrated software stack that can extract maximum performance, scale efficiently across systems, and operate reliably in production environments. With the latest ROCm 7.2 release, we're delivering a broad set of optimizations and software enhancements designed to improve developer productivity, runtime performance, and enterprise readiness.
In this blog, we highlight the latest ROCm 7.2 enhancements for AMD Instinct GPUs, designed to boost AI and HPC performance. Learn how hipBLASLt and GEMM optimizations, FP8/FP4 support in rocMLIR and MIGraphX, and topology-aware communication with GDA and RCCL deliver higher throughput and lower latency. We also cover AI model tuning for AMD Instinct MI300X and MI350 GPUs and Node Power Management (NPM) for efficient multi-GPU operation—together enabling faster, more scalable, and reliable AI workloads.