AI training workloads are pushing the limits of modern GPU architectures. With the release of AMD ROCm 7.0 software , AMD is raising the bar for high-performance training by delivering optimized support for LLM workloads across the JAX and PyTorch frameworks. The latest v25.9 Training Dockers demonstrate exceptional scaling efficiency for both single-node and multi-node setups, empowering researchers and developers to push model sizes and complexity further than ever.
By integrating Primus , a unified and flexible LLM training framework, the PyTorch training docker streamlines LLM development on AMD Instinct GPUs. Primus now supports both TorchTitan and Megatron-LM backends, while Primus-Turbo accelerates Transformer models, further boosting training throughput on AMD Instinct MI355X GPUs. Try Primus-Repo here .