HCCL is the Huawei Collective Communication Library. It is the part of Huawei's CANN software stack that moves tensors between Ascend NPUs during distributed training and inference, and it plays the same role on Ascend that NCCL plays on NVIDIA GPUs: every gradient all-reduce, every tensor-parallel all-gather and every expert-parallel all-to-all on an Ascend cluster goes through it. If you train on Ascend with PyTorch, you meet it as backend="hccl"; with MindSpore, as the default communication backend.
One naming trap first. Intel Gaudi also ships a library called HCCL, the Habana Collective Communications Library. It is a different product with a different API surface and different hardware underneath; for Gaudi see Intel Gaudi and GPU Max, in depth. This article is about Huawei's HCCL only.
The goal is practical: how HCCL bootstraps a job, which links it uses, how it picks an algorithm, what its buffers and timeouts cost, and how to measure and debug it. Variable names, ranges and defaults were checked against the CANN 8.x HCCL documentation on 2026-10-05; behaviour varies by release and Atlas product, so confirm them for yours.
Where HCCL sits
HCCL sits between the framework and the hardware. Above it, PyTorch reaches it through the torch_npu adapter, which registers the npu device and the hccl process-group backend, and MindSpore reaches it through its own communication module. You can also call the C API directly. Below it, three kinds of link carry traffic, and confusing them is the root of most early failures.
- Host NIC. An ordinary Ethernet interface on the server, used over TCP to exchange connection information when communicators are created. It carries almost no tensor data, but if HCCL picks the wrong one, nothing starts.
- Intra-server links. NPUs inside one server talk over Huawei's HCCS interconnect. HCCL chooses the intra-server algorithm (Mesh, Ring, Double-Ring or Star) automatically from the hardware topology; you cannot override it.
- Device NICs. On Atlas training servers each NPU has its own RoCE network port, with its own IP address. Inter-server collectives run NPU to NPU over these ports and do not touch the host CPU or host NIC. That is why a rank table lists a
device_ipfor every NPU.
Bootstrapping a communicator
A communicator is the set of ranks that take part in a collective, plus the connections and buffers between them. HCCL can build the world communicator in two ways, and frameworks use one or the other.
From a rank table. A JSON file describes every server, every device on it, each device's RoCE IP and its global rank. HcclCommInitClusterInfo(clusterInfo, rank, &comm) reads it, and MindSpore jobs point at it with the RANK_TABLE_FILE environment variable. The rank you pass must match the rank_id recorded for your device. An abridged single-server table looks like this:
{
"version": "1.0",
"server_count": "1",
"server_list": [
{
"server_id": "10.0.0.11",
"device": [
{"device_id": "0", "device_ip": "192.168.100.101", "rank_id": "0"},
{"device_id": "1", "device_ip": "192.168.101.101", "rank_id": "1"}
]
}
],
"status": "completed"
}From root info. Rank 0 generates a root handle, the handle is shared through some out-of-band channel, and every rank joins with it. This is the path PyTorch takes: the process group's TCP store plays the out-of-band role, so you set MASTER_ADDR and MASTER_PORT as you would for NCCL and no rank table is needed.
Either way, HCCL needs to know which host interface to use for the setup traffic. The documented order is: HCCL_IF_IP if set (one IPv4 or IPv6 address of a host NIC), then HCCL_SOCKET_IFNAME (interface names, with = for an exact match and ^= to exclude), then non-Docker, non-loopback interfaces in name order, then Docker interfaces, then loopback. On a host with a docker0 bridge and two data-centre NICs, set one of the two variables explicitly rather than relying on name order.
Hierarchy and algorithm choice
HCCL treats the cluster as a hierarchy: level 0 inside a server, level 1 between servers, and level 2 between supernodes on products that have them. Each level has its own bandwidth, so a collective is decomposed into per-level phases, for example a reduce-scatter inside each server, an all-reduce of the shards across servers, and an all-gather back inside each server. The inter-server phase is where algorithm choice matters most, and HCCL offers several:
| Algorithm (HCCL_ALGO name) | Steps | Good fit | Constraint |
|---|---|---|---|
| Ring (ring) | linear in servers | few servers, congested networks | latency grows with p |
| Recursive halving-doubling (H-D_R) | logarithmic | latency-bound, small to medium messages | requires a power-of-two server count |
| Nonuniform hierarchical ring (NHR) | logarithmic | many servers | none documented |
| Nonuniform Bruck (NB) | logarithmic | many servers | none documented |
| Pipeline | overlaps intra and inter links | large messages, many NPUs per server | needs enough data to fill the pipe |
| Pairwise | linear | AlltoAll family only | avoids one-to-many congestion |
HCCL picks among these by operator, message size and topology. To override the inter-server choice, set HCCL_ALGO="level0:NA;level1:NHR"; level 0 accepts only NA because intra-server selection is automatic. Which operators honour which level-1 values varies by product and release, so check the list for yours before pinning one.
The cost model behind the table is the usual alpha-beta one. With per-message latency alpha, per-byte time beta and p participants, a ring all-reduce of n bytes costs about 2(p-1)alpha + 2(p-1)/p n beta: bandwidth-optimal but with many latency terms. Halving-doubling costs about 2 log2(p) alpha plus the same bandwidth term, so it wins when messages are small and p is large. For the full derivation of these bounds and NCCL's equivalent choices, see NCCL Collectives.
Using HCCL from PyTorch
From PyTorch, HCCL looks like any other process-group backend. Importing torch_npu is what makes the npu device and the hccl backend available; the launch is the familiar torchrun. The script below creates the group, checks correctness with a known sum, then measures all-reduce bandwidth over a range of sizes.
import os, time
import torch
import torch_npu # registers torch.npu and the "hccl" backend
import torch.distributed as dist
def busbw(nbytes, seconds, world):
# bus bandwidth for all-reduce, comparable across world sizes
return nbytes / seconds * 2 * (world - 1) / world / 1e9
def main():
local = int(os.environ["LOCAL_RANK"])
torch.npu.set_device(local)
dist.init_process_group(backend="hccl")
rank, world = dist.get_rank(), dist.get_world_size()
# correctness first: sum of (rank + 1) over all ranks
t = torch.full((1024,), float(rank + 1), device="npu")
dist.all_reduce(t)
expected = world * (world + 1) / 2
assert torch.allclose(t, torch.full_like(t, expected)), t[:4]
for mib in (1, 16, 64, 256, 1024):
x = torch.ones(mib * 2**20 // 4, device="npu") # float32
for _ in range(5): # warm-up
dist.all_reduce(x)
torch.npu.synchronize()
t0 = time.perf_counter()
iters = 20
for _ in range(iters):
dist.all_reduce(x)
torch.npu.synchronize()
dt = (time.perf_counter() - t0) / iters
if rank == 0:
print(f"{mib:5d} MiB {dt*1e3:8.2f} ms busbw {busbw(x.numel()*4, dt, world):6.1f} GB/s")
dist.destroy_process_group()
if __name__ == "__main__":
main()
# torchrun --nnodes=4 --nproc-per-node=8 --rdzv-backend=c10d \
# --rdzv-endpoint=10.0.0.11:29500 allreduce_bench.pyTwo details matter. Call torch.npu.synchronize() around the timed loop, because collectives are queued on a stream and return before they finish. And always check the result before timing: a group that silently formed with the wrong ranks gives perfect bandwidth numbers for the wrong computation.
Buffers and the memory they cost
Each HCCL communicator reserves a shared buffer between peers, sized by HCCL_BUFFSIZE in MB. The default is 200 MB and the minimum is 1 MB. Data larger than the buffer is moved in pieces, so a bigger buffer can raise throughput for large messages, and a smaller one frees device memory.
The catch is that the cost is per communicator, and modern parallelism creates many of them. Take a rank in a job with data, tensor, pipeline and expert parallelism. It belongs to the world group, a data-parallel group, a tensor-parallel group, one or two pipeline peer groups and an expert-parallel group: six communicators is ordinary. At the default that is at least 1.2 GB of device memory per NPU before HCCL's other bookkeeping, which is a full layer of activations on some models. If you hit out-of-memory errors only at scale, or only when you add a parallelism dimension, count your communicators before you shrink the batch.
The tuning rule is simple. Raise the buffer for the few large, bandwidth-bound groups, such as data-parallel gradient buckets of hundreds of MB, and leave it at the default or lower it when memory is tight and messages are small, as in tensor-parallel activations. Measure both ways with the benchmark above; do not assume.
Worked example: reading a 32-rank benchmark
Suppose a job runs on 4 servers with 8 NPUs each, so 32 ranks, and the benchmark reports 95 ms for a 1 GiB float32 all-reduce. The algorithm bandwidth is 1,073,741,824 bytes / 0.095 s, about 11.3 GB/s. Bus bandwidth multiplies by 2(p-1)/p = 62/32, giving about 21.9 GB/s. That number is what you compare against hardware.
Inside a server, HCCS is far faster than that, so the inter-server phase is the bottleneck. In a hierarchical all-reduce each server first reduces to shards, then the 8 NPUs each move their shard across the network in parallel. The question is whether the per-NPU network traffic is near its port's line rate. For a 100 Gb/s port, line rate is 12.5 GB/s per direction; for 200 Gb/s, 25 GB/s. Read your own port speed and work out the inter-server share of the time from the per-level breakdown in the profiler, rather than assuming.
Now the decisions. If small messages, up to a few MiB, show poor bus bandwidth while large ones are fine, you are latency-bound: try HCCL_ALGO="level0:NA;level1:H-D_R" with 4 servers (a power of two) and compare. If large messages are poor, check that all 8 device NICs are up and on the right network, that the switch fabric is not oversubscribed (see Rail-Aligned Topology for why NIC-to-switch wiring matters), and try a larger HCCL_BUFFSIZE. Change one variable at a time and keep the table of results with the job config.
Failure modes and timeouts
Collective failures look alike from the outside: the job stops making progress. Two timeouts decide how long you wait before an error. HCCL_CONNECT_TIMEOUT limits socket setup between devices; it takes an integer in [120, 7200] seconds and defaults to 120. HCCL_EXEC_TIMEOUT limits how long a device waits for peers during execution; its range is (0, 17340] seconds and the default is 1836. That default means a single dead rank can freeze a whole cluster for about half an hour before anything reports an error, so set it to a few times your longest healthy step instead.
| Symptom | Likely cause | First check |
|---|---|---|
| Connect timeout at start-up | wrong host NIC chosen, firewall, or one rank started late | HCCL_IF_IP or HCCL_SOCKET_IFNAME on every rank; rank start times |
| Rank table rejected | rank_id or device_ip does not match the device | regenerate the table from the actual device NIC IPs |
| Hang mid-training, then exec timeout | a rank crashed or diverged and never entered the collective | first error in each rank's log, not the last |
| Hang only with some branches of code | ranks issue collectives in different orders | log the collective sequence per rank and diff |
| Device OOM only at scale | many communicators each holding HCCL_BUFFSIZE | count communicators per rank |
| Bandwidth drops on one server | a device NIC link is down or degraded | the vendor's device NIC tool (hccn_tool on Atlas hosts) |
The order bug deserves emphasis because it survives code review. Collectives are matched by position, not by name. If rank 3 skips a logging all-reduce because its batch was empty, rank 3's next gradient all-reduce pairs with everyone else's logging one, and the job either hangs or quietly averages the wrong tensors. Make every collective unconditional, or make the condition itself agreed by an all-reduce first.
Operational guidance
Running HCCL in production is mostly about making failures fast and legible.
- Pre-flight every job. Run the correctness check from the benchmark with the real world size before loading the model. It costs seconds and catches NIC, rank table and firewall errors before an hour of data loading.
- Pin the network explicitly. Set
HCCL_IF_IPorHCCL_SOCKET_IFNAMEin the launcher, not in a shell profile that some nodes lack. - Right-size timeouts. Keep
HCCL_CONNECT_TIMEOUTlong enough for the slowest node to start (graph compilation can be slow on first run) and cutHCCL_EXEC_TIMEOUTso a dead rank fails the job in minutes. Pair it with a restart from checkpoint. - Keep a baseline table. Record bus bandwidth at 1, 64 and 1024 MiB per world size after every driver, firmware or CANN upgrade. A 20 percent drop is a regression to explain, not noise.
- Overlap communication. Bucketed gradient all-reduce that runs while backward continues hides most data-parallel cost; the scheduling ideas carry over unchanged from Collective Communication Overlap.
Trade-offs
Porting from NCCL to HCCL is cheap at the framework level: the process-group API is the same. The costs are elsewhere: thinner tooling and community answers, documentation versioned per CANN release, and variables that differ between Atlas products. Pinning HCCL_ALGO can buy a few percent on a stable cluster and cost far more when the cluster size changes, so pin only with a measurement attached. A larger buffer trades device memory for bandwidth, and a short exec timeout trades false aborts during a slow checkpoint for fast recovery from real failures. For how intra-node interconnect speed shapes these choices on NVIDIA hardware, compare NVLink.
What to do next
- Confirm which HCCL you run (Huawei Ascend or Intel Gaudi) and which CANN release, and keep that release's environment-variable reference bookmarked.
- Set HCCL_IF_IP or HCCL_SOCKET_IFNAME explicitly in your launcher.
- Run the benchmark above at your production world size and save the bus bandwidth table.
- Lower HCCL_EXEC_TIMEOUT from the 1836 s default to a few times your step time.
- Count communicators per rank and budget HCCL_BUFFSIZE memory for each.
- Audit training code for conditional collectives and make them unconditional.
- Try one HCCL_ALGO level-1 override against the default, and keep it only if it wins.