Power & SitesUnited States

Growing Pains: How Distributed AI Training Changes The Network Between Datacenters

Credited to The Next Platform · nextplatform.com

Useful
Photograph · The Next Platform

Large-scale AI training has already escaped the confines of a single datacenter. Google said Gemini was trained synchronously across clusters in multiple locations; Microsoft has connected AI data centres in Wisconsin and Georgia into what it describes as one distributed AI supercomputer; AWS has connected AI compute clusters across wide areas to allow Anthropic to build Claude models; Meta has built high-capacity datacenter interconnects to support model training; and CoreWeave and Google Cloud recently announced cross-cloud training with a private interconnect, with Azure likely to follow later in the year.

This growing geographic spread of model training is partly due to the limits of single datacenters or campus clusters, which can be strained as they try to meet the compute and power demands of new models. Cisco estimates that training models today can require clusters with tens of thousands of GPUs. By 2030, the largest individual frontier training runs could draw 4-16GW of power, according to researchers at Epoch AI. Distributing compute allows hyperscalers and datacenter operators to build in locations with more available power or space and fewer planning constraints.

How To Train An LLM

At the core of an LLM is a neural network with billions of numerical parameters, or weights, that are adjusted during training. A batch of training data passes through the model, which makes a prediction. The system measures the error and calculates how the weights should change. The weights are updated and the process starts again. At the scale of modern models, this work is distributed across thousands of GPUs and other accelerators, such as AWS Trainium and Google TPUs.

Different GPUs can process separate batches of data or different parts of the model. But they cannot work entirely independently: at various points they must exchange results and synchronize before the next stage of training can begin.

What networks carry between the accelerators is generally not the text or images used to train the model, but large arrays of numerical data such as gradients and intermediate results. Thousands of accelerators may need to exchange and combine that data at roughly the same time.

"If you look at any training job, it is basically a repetition of compute and synchronization," explains Ramesh Sivakolundu of Cisco's Silicon One architecture team. "You compute the weights, communicate those weights to synchronize every GPU in the cluster, and then restart the computation. That means bursty traffic happens periodically throughout the training process."

Thousands of accelerators can finish a computation phase and start communicating at the same time. This can lead to "incast" problems where many senders converge on the same destination and traffic arrives faster than the next link can handle.

The fundamental issue lies with synchronization. In a synchronous training job, one group of GPUs cannot simply ignore a slower group and continue independently. The network-wide exchange must complete before the participating GPUs can move on to the next stage. A network delay therefore becomes compute delay.

"You don't want the network to be the bottleneck in the training process," says Sivakolundu. "The idea is to keep the GPUs fully occupied and not have the network introduce additional latency into the overall training program."

That does not mean that every instance of network congestion or packet loss brings the entire model training to a complete halt. However, a delayed or dropped flow may require retransmission, prolong a collective communication operation, and leave other accelerators waiting. More serious failures can require training to return to a saved checkpoint.

Different Kind Of Interconnect

"Inside a single datacenter, you can assume you have built a full network connecting the GPU racks with the bandwidth you planned for," says Itamar Gold, director of product management at Cisco. Gold’s point is that with scale-across, coordinated AI workloads may have to traverse a more constrained inter-site fabric than the network within each facility. Instead, he says, traffic destined for the other facility may be funneled through what is effectively a narrower pipe. A synchronized burst can turn that pipe into a temporary choke point.

Traditional DCI connects otherwise independent datacenter environments, and is typically built for redundancy, reach and workload distribution. The traffic that runs over them is often replication traffic or application data, and the flows are asynchronous, meaning one flow does not have to finish in lockstep with the others.

In scale-across AI training, the inter-site network becomes part of a coordinated distributed computation rather than simply moving data between otherwise independent systems. It is therefore essential to have sufficient bandwidth and predictable delivery.

From Rack To Region

Within a rack, bandwidth requirements can be enormous – up to 800 Gb/sec and 1.6 Tb/sec – but the distances are short. High-speed serial lanes combine to make a GPU cluster appear as a single compute unit. When the AI fabric is extended across an entire facility, “scale out” racks will still require connections up to 1.6 Tb/sec. At these data rates, copper is generally confined to short reaches. Connections across rack rows and datacenter floors instead use pluggable optics, with scale-out links extending from hundreds of meters to around 2km.

“Scale-across” extends a coordinated AI fabric across multiple facilities, making inter-site network behavior part of workload performance. Depending on the architecture, those links may span metro or regional distances, and potentially farther depending on the architecture. According to Cisco modelling of large distributed AI deployments, aggregate bandwidth requirements can reach roughly 14x a conventional DCI baseline.

For starters, short-range datacenter optics are not designed to carry 400-800 Gbps signals over hundreds of kilometers. For these distances, operators need coherent optics that use sophisticated modulation and high-performance digital signal processing to keep high-rate optical signals usable over metro and regional fiber spans, compensating for physical impairments that become more pronounced with distance. The number of ports matters too. Cisco estimates that connecting two 100 megawatt AI sites could require 12,000 to 32,000 coherent optical ports at the scale-across layer, compared with roughly 1,000 to 2,000 for conventional DCI between comparable facilities. At 400 Gb/sec or 800 Gb/sec per port, that adds up to multi-petabit aggregate capacity. Moving that much data is also a power and space problem. Coherent pluggables put the coherent transponder function directly into router or switch ports, avoiding a separate bank of standalone DWDM transponders. Cisco says this reduces the additional rack space, power and cooling needed for the inter-site optical layer which is an important consideration when the AI factories themselves are already power-constrained.

The Latency Problem

Even with sufficient bandwidth and suitable optics, distance adds latency. If congestion develops, the network needs time to signal the sender to slow down. Over a local AI fabric, feedback can arrive quickly. However, if the connection is stretched across 100 kilometers of fiber, considerably more data can be transmitted before the sender even learns that there is a problem.

Cisco calculates that on an 800 Gb/sec connection spanning 100 kilometers, around 100 MB of data can already be in transit during that feedback interval. Cisco argues that these longer feedback loops increase the value of networking silicon designed with substantially deeper buffering. The shallow-buffer switching architectures designed for very low latency inside a datacenter have less capacity to absorb traffic while congestion feedback crosses a long-distance link.

"The reason you need deeper buffers as the distances get longer is that you don't get feedback about the transfer of packets for a longer period of time," says Sivakolundu. "If you don't buffer deeply enough, you can end up dropping and retransmitting packets, increasing latency."

Deep buffering absorbs a temporary burst or disruption while traffic management and congestion signaling address the underlying problem. Sustained congestion still needs more capacity or a change in how traffic is distributed.

The same buffer also helps with oversubscription. If the combined capacity feeding into an inter-site connection is greater than the capacity of the link itself, a synchronized burst can arrive faster than the link can drain it. Buffering gives that excess traffic somewhere to wait without getting dropped while congestion controls respond.

Rather than dividing the available buffer into fixed amounts for each port, a shared buffer is a common pool of packet memory that can be allocated wherever congestion appears. That lets a busy link absorb a larger burst without dropping packets.

"The longer you're reaching out, the more you may need to buffer," says Gold. "You also need the flexibility to put as much buffer as possible where it is needed, perhaps on one port or a few ports where you've identified a problematic flow. Our approach is a single, fully shared buffer."

AI training does offer one useful characteristic: much of its collective communication is repetitive and therefore more predictable than DCI traffic. Cisco's approach is to use proactive congestion management to steer or schedule traffic before a predictable burst creates a problem, while retaining deep buffering as a safety net for transient congestion and unpredictable events such as a link failure. The two approaches are complementary: one tries to avoid the queue while the other gives the network somewhere to put the traffic when a queue forms.

Cisco’s Bet On Co-Design

Cisco’s scale-across architecture uses the Silicon One P200, a 51.2 Tb/sec programmable, deep-buffer routing processor included in the Cisco 8223 and Cisco N9000 platforms. Those systems can be paired with Cisco's 400G and 800G coherent pluggable optics for the long-distance links, while open line systems such as Cisco Open Transport 3000 and Cisco NCS 1014 handle the optical transport layer where required.

Coherent pluggables generate the DWDM wavelengths directly from the routing system, while the line system amplifies and carries those wavelengths across the fiber plant. For links requiring multiple parallel fiber pairs, Cisco's Open Transport 3000 uses a multi-rail design that combines the optical components for multiple rails onto one line card. Cisco claims this reduces power per rail by 75 percent and rack space by 80 percent. Where a separate transport system is required, the NCS 1014 can provide 12.8 Tb/sec of capacity from a 1RU line card.

The P200 has a programmable run-to-completion network processor and P4 tooling, which Cisco says allows protocol support, telemetry and other packet-processing features to be developed in software rather than making every change dependent on a new silicon generation.

For scale-across, where operators are still testing different topologies and traffic-management methods, "the same programmable flexibility can let network functions evolve without every new requirement forcing a chip replacement,” says Sivakolundu.

Cisco also includes hardware-based protection for inter-site traffic in that design, its argument being that encryption should not create another processing bottleneck for the training fabric.

A key part of Cisco's pitch is the co-design. A networking ASIC is defined years before the finished router reaches a datacenter. Cisco argues that because its silicon, system and optics teams sit within the same company, requirements for port density, buffering, power, telemetry and the optical interfaces can feed back into the chip design rather than the system team receiving a commercial ASIC and working out what it can build around it.

That does not mean scale-across has settled on a single approved architecture. Cisco itself divides the problem into campus, metro, and regional deployments, and different operators are experimenting with various combinations of networking and training techniques.

Implementations vary significantly by operator, including distance, topology and workload design, while the common problem is maintaining coordinated AI performance across sites. The technology is still at an early stage, however. "We are seeing early deployments now, but it is a process and I think it will take a few years," says Gold.

Designing a network for distributed training will not eliminate the latency of distance, but it can keep that interconnect from becoming the bottleneck. That gives hyperscalers and datacenter operators the flexibility to add compute where power and space are available, while still treating multiple sites as part of the same training infrastructure.

Original · The Next Platform

FrontMethod