SoftBank and AMD have started joint validation of AMD Instinct GPUs for next-generation AI infrastructure by focusing on a GPU partitioning mechanism that lets a single GPU handle multiple AI workloads simultaneously. SoftBank developed an Orchestrator system that divides AMD Instinct GPU resources based on workload requirements: model size, number of concurrent executions, and memory needs. The system splits compute workloads across multiple GPU instances running on individual Accelerator Complex Dies (XCDs). These can range from single instance mode (SPX mode) up to eight instances (CPX mode). The HBM memory pools are also divided into individual regions for each GPU instance to prevent latency spikes. The goal is to avoid the inefficiency of uniform GPU resource allocation, which can lead to GPU resource shortages or waste depending on workload demands.
SoftBank says the enhanced Orchestrator runs multiple AI applications on a single GPU with minimal resource strain. No performance figures have been shared yet, though SoftBank mentions improved resource allocation for small and mid-size language model workloads. The company also plans to explore similar orchestration for other AI accelerators beyond AMD. A live demonstration is planned at the AMD booth during MWC Barcelona 2026 in March 2-5. SoftBank also published technical details on architecture and Orchestrator management methods on its Research Institute of Advanced Technology blog. In a recent report , we noted that AMD's next-gen Instinct MI455X accelerators, set to compete with NVIDIA's Vera Rubin, are running into serious manufacturing problems that are pushing back AMD's roadmap. Only limited production is expected this year, with mass production pushed to Q2 2027.