AWS Releases Trainium3 ASIC to Ease Reliance on NVIDIA Hardware

During its re:Invent conference in Las Vegas, AWS introduced its latest ASIC chip, Trainium3, designed for internal AI workloads and select external customers. This chip delivers 2.52 PetaFLOPS of FP8 compute per chip and increases on-chip memory capacity to 144 GB of HBM3E, with a memory bandwidth of 4.9 TB/s. Trainium3 supports both dense and expert-parallel model topologies and introduces compact data types, MXFP8 and MXFP4, enhancing the balance between memory and compute for real-time, multimodal, and long-context reasoning tasks. The chip is manufactured using TSMC's N3 3 nm node and is now available in Amazon EC2 Trn3 UltraServer instances.

Trn3 UltraServers can scale up to 144 Trainium3 chips in a single server, achieving approximately 362 FP8 PetaFLOPS. These servers can be combined into EC2 UltraClusters 3.0 for larger deployments. A fully equipped UltraServer offers about 20.7 TB of HBM3e memory and around 706 TB/s of aggregate memory bandwidth. It also features the NeuronSwitch-v1 fabric, which doubles the interchip interconnect bandwidth compared to the previous UltraServer. AWS reports significant generational improvements, with up to 4.4x higher performance, 3.9x greater memory bandwidth, and about 4x better performance per watt compared to the Trainium2. Additionally, there are notable enhancements in inference and token efficiency for various Amazon services.
Read full story