Exploring the breakthrough innovations shaping our world. From AI infrastructure and robotics to biotech, quantum computing, and spatial tech.
As AI training systems grow from thousands of accelerators to tens of thousands and, in some cases, 100,000 or more, the networking problem changes substantially. Adding more GPUs or TPUs does not automatically produce proportional gains in training performance. At this scale, the network connecting those processors can become just as important as the processors themselves.
A modern AI supercluster is not simply a very large collection of servers. It is a distributed computing system in which thousands of accelerators repeatedly exchange data while training a single model. Every additional switch, cable, link, and routing decision can affect how efficiently that system operates.
The challenge is particularly visible during collective communication. Training workloads use techniques such as data parallelism, tensor parallelism, and pipeline parallelism to distribute computation across many accelerators. These approaches reduce the computational and memory burden on individual devices, but they also increase the amount of information that must move between them.
For a 100,000-accelerator system, the network therefore has to deliver enormous aggregate bandwidth while keeping latency, congestion, and failure recovery under control. The goal is not simply to connect every machine. It is to make the entire fabric behave as predictably as possible under sustained, highly synchronized traffic.
Large AI models are commonly trained using multiple forms of parallelism. Data parallelism distributes batches across different accelerator groups, while tensor and pipeline parallelism divide parts of the model or its execution across devices.
Each approach creates communication requirements of its own. Some exchanges occur frequently between closely related devices, while others involve much larger groups of accelerators.
One of the best-known collective operations is All-Reduce. During an All-Reduce, every participating accelerator contributes data to a reduction operation, and the resulting aggregate is made available to all participants. The operation is essential for synchronizing information such as gradients in many distributed training workloads.
The important point is that All-Reduce does not mean every GPU must establish a direct point-to-point connection with every other GPU. Instead, software and hardware use communication algorithms and network paths to perform the collective efficiently. Ring, tree, hierarchical, and topology-aware approaches are among the strategies that can be used depending on the system architecture.
At very large scale, however, even efficient collective algorithms can generate enormous bursts of east-west traffic. If the network cannot absorb those bursts, queues grow, packets may be retransmitted or delayed, and some accelerators finish their local computation before the required communication has completed.
That creates a straggler effect. A large portion of the cluster can be ready for the next stage of training while waiting for a smaller portion of the communication operation to finish.
For expensive accelerator fleets, that idle time matters. A network that looks fast in a simple bandwidth test can still become a performance bottleneck when thousands of devices communicate simultaneously.

Conventional enterprise data centers have historically used hierarchical designs built around access, aggregation, and core layers. These architectures work well for many applications in which traffic primarily moves between end users and centralized services.
AI training has a different traffic pattern.
Instead of mostly north-south communication between clients and servers, large distributed training jobs generate substantial east-west traffic between peer accelerators. GPUs may exchange information with other GPUs throughout the cluster during successive stages of computation and synchronization.
This changes the requirements placed on the network.
A highly oversubscribed hierarchy can create congestion as traffic moves through shared uplinks. Even when individual links have high capacity, the aggregate bandwidth available between different portions of the cluster may be insufficient for a demanding collective operation.
This is where bisection bandwidth becomes important. In simplified terms, bisection bandwidth describes how much traffic the network can carry when the system is divided into two sections and communication must cross between them. A fabric with insufficient bisection bandwidth can become congested even when many individual links are operating below their maximum line rate.
For AI superclusters, network designers therefore pay close attention to topology, switch radix, link speeds, oversubscription, path diversity, and the placement of accelerators within the fabric.
Multi-stage Clos networks, often associated with Fat-Tree architectures, are among the most important approaches for building large-scale data center and AI networks.
A Clos fabric uses multiple stages of switches rather than relying on a single large central switch. Leaf switches connect to servers or accelerator systems, while higher-level switches provide paths between different sections of the cluster.
A carefully engineered Clos design can provide a large number of parallel paths between endpoints. This makes it possible to distribute traffic across multiple links and avoid relying on a single physical path for large collective operations.
One of the major advantages of Clos architectures is scalability. Additional switches and links can be introduced in a structured way as the cluster grows.
With appropriate link and port provisioning, the fabric can provide high aggregate bandwidth and relatively predictable communication performance. The architecture is also well suited to multipath routing, allowing traffic to use different paths through the network.
The trade-off is physical complexity.
A very large Clos network can require substantial numbers of switches, optical transceivers, cables, power connections, and rack-space resources. As the number of accelerators increases, the amount of infrastructure required to maintain high connectivity grows quickly.
Cabling is not merely an installation problem. Large optical fabrics also affect power consumption, cooling requirements, maintenance procedures, and the physical layout of the data center.
For a 100,000-plus accelerator system, these considerations become part of the architecture rather than an afterthought.

Another approach is to reduce the amount of centralized switching by using direct network topologies.
High-performance computing has long explored designs such as torus and Dragonfly networks. These architectures connect computing nodes or groups of nodes more directly and can reduce the amount of switching infrastructure required for certain system configurations.
A Dragonfly network organizes nodes into groups. Local connections provide communication within a group, while higher-bandwidth global links connect different groups.
This structure can provide relatively short paths between many endpoints while reducing the number of traditional multi-stage switch layers.
The appeal becomes clearer at large scale. A network that relies on fewer layers of switching may reduce cabling and switch-port requirements compared with an equivalently provisioned large Clos fabric.
There is a trade-off, though.
Global links can become valuable resources when many workloads simultaneously send traffic between groups. Effective routing and congestion-management strategies are therefore essential. The network must distribute traffic intelligently rather than allowing a small number of links to become persistent hot spots.
Torus designs connect nodes in multiple dimensions, creating regular paths through the network. They have been used extensively in HPC systems because their structure is predictable and can map well to workloads with particular communication patterns.
The limitation is that topology efficiency depends heavily on workload behavior. A topology that performs well for a communication pattern with strong locality may be less attractive for workloads that require broad, irregular communication across the entire system.
For AI clusters, this is one reason topology cannot be evaluated independently from the training software and collective communication algorithms.

Topology determines how devices are connected, but the transport technology determines how data moves across those connections.
RDMA, or Remote Direct Memory Access, is particularly important in high-performance AI and HPC environments. RDMA allows data to move between hosts or accelerators while minimizing involvement from the host CPU and operating system.
This can reduce software overhead and help deliver low-latency communication for workloads that exchange large quantities of data.
Two major approaches are commonly discussed in modern AI infrastructure: InfiniBand and RoCE.
InfiniBand is a high-performance networking architecture designed for demanding computing environments. It provides mechanisms for flow control, congestion management, and low-latency communication.
Its tightly integrated ecosystem has made it widely used in HPC systems and large AI deployments where predictable high-performance communication is a major requirement.
The key advantage is not simply raw link speed. The complete stack, including adapters, switches, routing, communication libraries, and software support, is designed around high-performance distributed computing.
RoCE, or RDMA over Converged Ethernet, brings RDMA capabilities to Ethernet networks.
The approach is attractive because Ethernet infrastructure is widely deployed across data centers and can provide a broad hardware ecosystem. Modern RoCEv2 implementations can use mechanisms such as Priority Flow Control and Explicit Congestion Notification to manage traffic in demanding environments.
The challenge is configuration.
A large RoCE deployment requires careful attention to congestion behavior, queue management, traffic classes, routing, and failure scenarios. Ethernet provides flexibility, but achieving stable performance for highly synchronized AI workloads requires substantial engineering.
One of the most important lessons from large AI systems is that network hardware cannot be designed independently of the software stack.
A topology may look excellent in a generic network benchmark but behave differently when thousands of accelerators execute the same collective operation at nearly the same time.
Training frameworks and communication libraries can take topology into account when selecting communication algorithms. If certain GPUs are connected through shorter or higher-bandwidth paths, the software can sometimes organize communication to take advantage of that structure.
This creates a feedback loop between hardware and software.
Network topology influences the best collective algorithm, while the collective communication pattern influences which topology provides the best practical performance.
For very large clusters, optimizing this interaction can be more valuable than simply increasing the speed of an individual link.
Large systems also have a different relationship with hardware failures.
When thousands of components operate continuously, failures are no longer unusual events. Power supplies, optical transceivers, cables, network ports, memory components, and accelerator devices can all experience faults over the lifetime of a system.
The network therefore needs redundancy.
Multiple paths can allow traffic to continue when an individual link or switch becomes unavailable. Routing systems can redirect traffic around failed components, while higher-level software can detect degraded nodes and adjust workload placement.
The exact recovery time depends on the failure type, detection mechanism, routing design, transport protocol, and application behavior. The objective is not necessarily to make every failure invisible. Instead, the system should limit disruption and avoid turning a localized hardware problem into a cluster-wide training failure.
Training systems can also use checkpointing and fault-tolerant scheduling to reduce the cost of failures that cannot be completely hidden by the network.
At this scale, network resilience and application resilience become closely connected.

There is no single network topology that is ideal for every AI supercluster.
A large Clos fabric can provide strong connectivity and many parallel paths, but it may require substantial switch, transceiver, and cabling resources.
Direct architectures such as Dragonfly can reduce some infrastructure requirements, but they place greater importance on routing, traffic management, and workload-aware design.
InfiniBand can provide an integrated high-performance networking environment, while RoCE can leverage the broader Ethernet ecosystem. Neither choice eliminates the need for careful engineering.
The right architecture depends on several factors:
accelerator count and accelerator-to-network bandwidth;
expected communication patterns;
collective operations used by the training framework;
required bisection bandwidth;
acceptable oversubscription;
switch radix and port availability;
optical and cable density;
power and cooling constraints;
fault tolerance requirements;
routing and congestion-control capabilities; and
the software stack running on the cluster.
These variables need to be considered together rather than optimized independently.
Building an AI supercluster with 100,000 or more accelerators is not simply a matter of installing more processors. At this scale, the network becomes part of the computing architecture itself.
Distributed training creates large volumes of synchronized communication, and the efficiency of that communication depends on topology, bandwidth, latency, routing, congestion management, and fault tolerance. A cluster with exceptionally fast accelerators can still waste substantial compute capacity if its communication fabric cannot keep those devices supplied with data.
Multi-stage Clos networks remain an important option for highly connected AI fabrics, while direct approaches such as Dragonfly and torus architectures offer different trade-offs in switching, cabling, routing, and scalability. InfiniBand and RoCE provide complementary approaches to high-performance RDMA networking, with each requiring careful system-level engineering.
The most effective design is therefore not defined by a single topology or protocol. It comes from matching the network to the communication behavior of the workloads, provisioning enough aggregate bandwidth, managing congestion intelligently, and building sufficient redundancy to keep large distributed jobs running when individual components fail.
At 100,000-accelerator scale, the network is no longer just the infrastructure connecting the computers. It is one of the systems that determines how effectively the computers can operate as a single machine.