Data center networking is one of the most overlooked factors by the public, which is in fact responsible for all communication between nodes. However, NVIDIA knows that data centers with millions of GPUs are on the horizon, and for the fastest AI models, they need to be interconnected, even across multiple facilites. That is why NVIDIA today unveiled Spectrum-XGS Ethernet, an extension of the Spectrum-X networking platform engineered to interconnect multiple, geographically separated data centers into a single, giga-scale AI super-factory . The company states that Spectrum-XGS eliminates the capacity limits of single facilities by introducing distance-aware networking, which delivers predictable, low-latency performance across campuses, cities, and continents.
The technology is mainly delivered through software and firmware updates to existing Spectrum-X switches and ConnectX SuperNICs rather than new silicon. Spectrum-XGS offers auto-adjusted congestion control optimized for long-haul links, precise latency management to minimize jitter, and comprehensive end-to-end telemetry, which allows operators to visualize and control network traffic across multiple sites. NVIDIA reports that these changes nearly double the NCCL (Collective Communications Library) throughput for multi-GPU, multi-node training jobs and large-scale experiments, thereby improving efficiency for distributed AI workloads. NVIDIA framed Spectrum-XGS as a new axis of growth for AI infrastructure: after scaling up inside servers and scaling out inside data centers, scale-across connects facilities into unified compute fabrics.