Open Source LLM

(PR) AMD Details Single Node & Distributed Inference Performance on Instinct MI355X GPU

The rapid growth of GenAI and LLM workloads, from agentic workflows and multi-step tool use to retrieval-augmented reasoning, demands inference infrastructure that is fast, adaptable, and highly optimized. While AMD continues to invest in general-purpose inference frameworks such as vLLM and SGLang, ATOM provides the most direct path to peak Instinct MI355X GPU performance for modern reasoning and MoE-heavy workloads, which increasingly dominate frontier LLM architectures.

Over the past months, AMD have implemented numerous optimizations to improve both single-node performance and multi-node distributed inference for DeepSeek-R1 on the MI355X GPU.
Read full story