A dragonfly is a two-level network built for one purpose: to connect a very large number of endpoints with few long cables and a small hop count. Routers are grouped. Inside a group every router links to every other, so the group behaves like one very high-radix virtual router. Groups then link to each other, ideally every group to every other group, over a modest number of long global links. Any endpoint can reach any other in at most three router-to-router hops: one local hop to the router that owns the right global link, the global hop, and one local hop to the destination router.

The design comes from a 2008 paper by Kim, Dally, Scott and Abts, and it now runs most of the largest supercomputers through the Cray Aries and HPE Slingshot interconnects. For anyone running training jobs on such a machine, the topology is not background detail. Its cheap global tier is also its scarce tier, and whether an all-reduce or an all-to-all runs at full speed depends on how many global links your traffic funnels through and on which routing mode the fabric picks. This article builds the topology from its parameters, shows the traffic pattern that breaks it, explains the routing that rescues it, and turns that into placement rules. For the tree it is usually compared with, see fat-tree topology.

The three parameters

Three numbers define a dragonfly. Each router has p ports to endpoints (NICs), a - 1 local ports to the other routers in its group (there are a routers per group), and h global ports to other groups. Router radix is therefore k = p + (a - 1) + h. A group has a*h global ports in total, so with one link between every pair of groups the network can hold at most g = a*h + 1 groups and N = a*p*g endpoints.

The original paper recommends a balanced design with a = 2p = 2h. The reasoning is about load. Under uniform random traffic, most packets leave their group, so each router must push roughly p endpoints' worth of traffic over its h global links, which argues for h near p. Each global packet also crosses up to two local hops, which argues for more local than global capacity, hence a about 2h. The calculator below does the arithmetic and also reports how many parallel global links each pair of groups gets when you build fewer than the maximum number of groups, which is how real machines are built.

def dragonfly(p, a, h, groups=None):
    """Size a dragonfly. p endpoints, a routers per group, h global ports per router."""
    radix = p + (a - 1) + h
    g_max = a * h + 1
    g = groups or g_max
    if g > g_max:
        raise ValueError(f"at most {g_max} groups with one link per group pair")
    global_ports = a * h                      # per group
    links_per_pair = global_ports // (g - 1)  # parallel global links between two groups
    return dict(radix=radix, groups=g, endpoints=a * p * g,
                endpoints_per_group=a * p, links_per_pair=links_per_pair)

print(dragonfly(p=4, a=8, h=4))
# {'radix': 15, 'groups': 33, 'endpoints': 1056, 'endpoints_per_group': 32, 'links_per_pair': 1}
print(dragonfly(p=16, a=32, h=17))
# {'radix': 64, 'groups': 545, 'endpoints': 279040, 'endpoints_per_group': 512, 'links_per_pair': 1}
print(dragonfly(p=16, a=32, h=17, groups=74))
# {'radix': 64, 'groups': 74, 'endpoints': 37888, 'endpoints_per_group': 512, 'links_per_pair': 7}

The second line is the published Slingshot shape. The Rosetta switch has 64 ports at 200 Gb/s; a group is 32 switches fully connected with 31 ports each, 16 ports go to endpoints and 17 to global links. That gives the 545-group, 279,040-endpoint ceiling HPE quotes, with a switch-to-switch diameter of three. Frontier was reported as 74 groups of 32 switches, so its groups are joined by bundles of about seven global links rather than one, which matters a great deal for the traffic analysis that follows. Note that Slingshot is close to, not exactly, balanced: h is 17 rather than 16.

Dragonfly: routers form fully connected groups; groups connect all-to-allGroup 0R0R1R2R3Group 1R0R1R2R3Group 2R0R1R2R3global linkglobal linkglobal linkMinimal path: local, global, localat most 3 router-to-router hopsValiant path: via a random third groupup to 5 hops, two global linksGrey: local all-to-all links inside a group. Amber: global links between groups (often optical).
Three groups of four routers. Local links (grey) make each group a clique; global links (amber) join groups. A minimal route uses at most one global link; a Valiant route detours through a third group and uses two.

Minimal routing: local, global, local

Minimal routing is mechanical. A packet from router R in group G to a router in group H first checks whether R itself owns a global link to H. If not, it takes one local hop to the router in G that does, crosses the global link, lands on some router in H and, if that is not the destination router, takes one more local hop. Local, global, local: three hops worst case, and every endpoint pair has a short path.

Compare this with a three-tier fat tree, where a packet that leaves its pod climbs to the core and back down for up to four switch-to-switch hops, but where the core is provisioned with as much bandwidth as the edge. The dragonfly wins on cables and switches: global links are the expensive optical ones, and a dragonfly needs roughly one global port per endpoint rather than several layers of them. It pays for that with non-uniformity. Bandwidth between two specific groups is only as wide as the bundle of global links joining them.

The traffic that breaks it

Uniform random traffic is the friendly case: each source picks a random destination, load spreads over all global links, and minimal routing gets close to full throughput. The hostile case is a group shift: every endpoint in group i sends to an endpoint in group i + 1. Under minimal routing, all of group i's traffic must cross the links joining i to i + 1 and nothing else.

Put numbers on it. In the 545-group Slingshot shape, a group holds 512 endpoints and there is one global link per group pair. A group shift squeezes 512 endpoints of traffic through one 200 Gb/s link: each endpoint gets about 1/512 of its injection bandwidth. In the 74-group build, seven parallel links share the load, which is still roughly 73 endpoints per link. The training analogue is not exotic. A pipeline-parallel job placed so that stage s lives in group s, with every rank in a stage sending activations to the matching rank in the next stage, is a group shift. So is a data-parallel ring whose consecutive members sit in consecutive groups and whose ring segments all start at once.

Valiant and UGAL routing

The standard fix is Valiant routing: send each packet first to a randomly chosen intermediate group, then minimally to its destination. Any traffic pattern becomes two phases of uniform random traffic, so no pattern can concentrate load on one bundle. The price is that every packet now uses two global links and up to five router hops, which halves the usable global bandwidth for patterns that minimal routing would have served perfectly, and adds latency.

Production fabrics therefore route adaptively. UGAL (universal globally adaptive load balancing), from the same research line, compares the minimal path with one randomly chosen Valiant path at the moment of injection and picks the one with the smaller product of queue occupancy and hop count. The local variant uses only the queues visible at the source router, which is cheap but sees congestion late; Slingshot and Aries add congestion information propagated from other switches. In sketch form:

def choose_route(src_router, dst_group, rng, bias=0):
    """UGAL-L style decision at the source router."""
    minimal = minimal_path(src_router, dst_group)            # <= 3 hops, 1 global link
    mid = rng.choice([g for g in all_groups()
                      if g not in (src_router.group, dst_group)])
    detour = valiant_path(src_router, mid, dst_group)        # <= 5 hops, 2 global links

    q_min = src_router.queue_depth(minimal.first_port)
    q_val = src_router.queue_depth(detour.first_port)
    # Weight queue depth by path length: a long path consumes more total capacity.
    if q_min * len(minimal.hops) <= q_val * len(detour.hops) + bias:
        return minimal
    return detour

The bias term is a tuning knob: a positive value favours minimal paths, which suits latency-sensitive small messages; a negative value favours spreading, which suits bulk collectives. Slingshot exposes per-traffic-class choices along these lines; the knobs and defaults differ by system release, so read your site's documentation.

Deadlock and virtual channels

Adaptive routes create cycles of buffer dependencies: a local link can carry traffic that is both just starting and just finishing its trip, and a full buffer at one router can wait on a full buffer at another that waits on the first. Dragonflies break the cycle with virtual channels. Each hop class gets its own buffer pool, and a packet may only move to a higher-numbered channel: for example local hops before the first global link use VC0, the first global hop VC1, local hops in the intermediate group VC2, and so on. Because channel numbers only rise along a path, no waiting cycle can close. Valiant paths need more channels than minimal ones, which is part of why router designers count buffers carefully.

What training traffic sees

Training traffic has three shapes, and each meets the dragonfly differently.

  • Ring and tree all-reduce for data parallelism. Bandwidth-bound and long-lived. If ring neighbours share a group, most of the ring stays on local links and only a few segments cross groups. If the scheduler scatters ranks across many groups, every segment is global and you have built a group shift.
  • All-to-all for expert parallelism and sequence or context parallelism. Close to uniform random traffic, which minimal routing handles well, but every byte that leaves a group competes for the global tier. Size the expert-parallel group to fit inside one network group when you can.
  • Point-to-point pipeline traffic. Small in total but latency-critical. Keep adjacent stages in one group or accept the extra hop; avoid placing stage s in group s for every s.

The pattern is the same as the hierarchy advice for any cluster: keep the heaviest, most frequent collective on the fastest, most local links. On a GPU system the first tier is the NVLink domain inside a node; the dragonfly group is the second tier, and the global links are the third. Rail-aligned designs solve the same problem on fat trees, and in-network reductions shrink the bytes that cross the scarce tier at all.

Worked example: a 2,048-endpoint job

Take a job of 2,048 endpoints on the 74-group Slingshot shape: 512 endpoints per group, seven global links per group pair. The job fits in four groups. Run it as 4-way data parallel, each replica a 512-rank block of tensor, pipeline and expert parallelism placed inside one group. The tensor and expert traffic, which is the heavy and frequent part, then stays entirely on local links. Only the data-parallel gradient all-reduce crosses groups, and a ring over four groups crosses only four group boundaries, each served by a bundle of seven links.

Now let the scheduler place the same job on 2,048 endpoints scattered across 40 groups, about 51 per group. Every 512-rank block spans about ten groups, and the all-to-all, which was local, now crosses the global tier. Adaptive routing keeps the fabric from collapsing, but the job shares global bundles with every other job on the machine, and step time becomes a function of the neighbours' traffic. The measurement that exposes this is simple: run the same all-to-all benchmark at both placements and compare bus bandwidth.

Failure modes

  • Placement fragmentation. A busy machine hands out whatever endpoints are free, so a job lands in many groups. Symptom: step time varies with the time of day. Fix: request group-aligned allocations and log the group of every rank at start-up.
  • Synchronised group shifts. Pipeline or ring layouts that map rank order onto group order. Symptom: one collective is much slower than the bandwidth model says, and global-link counters show a few bundles at saturation. Fix: permute the rank order or let adaptive routing spread the load.
  • Routing mode forced to minimal. Someone set a latency-optimised mode for one benchmark and left it. Symptom: adversarial patterns collapse. Fix: check the environment the job actually inherits.
  • Degraded global links. A failed optical link shrinks a bundle, and with one link per pair it removes the direct path. Routing survives through detours, but bandwidth between those two groups drops. Check fabric health before blaming the model.

Trade-offs against a fat tree

PropertyDragonflyNon-blocking fat tree
Diameter (switch hops)3 minimal, up to 5 with ValiantUp to 4 in three tiers
Long cablesAbout one global port per endpoint, few layersSeveral optical layers
Worst-case bandwidthLow under minimal routing; about half with ValiantFull bisection
RoutingAdaptive routing is essentialECMP works, adaptive helps
Placement sensitivityHigh: group boundaries matterModerate: pods matter
Typical homeLarge HPC systems (Aries, Slingshot)Most GPU training clusters

The summary: a dragonfly buys scale and cost efficiency with a scarce global tier and a dependence on good routing and placement. A fat tree buys predictability with more switches and optics. Neither is better in general; the question is whether your workload's heavy traffic can be kept inside groups. See GPU pod networks for how scale-up and scale-out tiers are mapped to parallelism in a GPU cluster.

What to do next

  1. Find your system's p, a and h, and run the calculator above to get endpoints per group and links per group pair.
  2. Log the network group of every rank at job start; it is usually derivable from the hostname or a site tool.
  3. Map the most bandwidth-hungry collective (tensor or expert parallel) inside one group before anything else.
  4. Benchmark all-reduce and all-to-all at your real placement, not on an idle reservation.
  5. Check which routing mode your MPI or NCCL transport inherits, and do not force minimal routing globally.
  6. Ask your site for group-aligned allocation and use it for jobs larger than one group.
  7. Re-measure after any fabric maintenance; a degraded global bundle changes the numbers.
Key takeaway: A dragonfly joins fully connected groups of routers with a thin layer of global links, giving a three-hop diameter at low cable cost. Its weakness is that tier: a group-shift pattern can funnel a whole group through one bundle, which adaptive routing such as UGAL mitigates by detouring through random groups at a bandwidth cost. Keep heavy collectives inside a group, keep rank order from mapping onto group order, and measure at your real placement.