The Graphcore Intelligence Processing Unit is the clearest example of an AI chip that does not try to be a better GPU. It removes the cache hierarchy and the large off-chip DRAM, and instead splits the die into 1,472 small processors, each with its own private SRAM. A compiler decides at build time where every tensor lives and when every byte moves. That choice makes some workloads very fast and makes others awkward or impossible, and almost every practical question about the IPU follows from it.

This article explains the second-generation GC200 chip from first principles: the tile, bulk synchronous execution, how the Poplar graph compiler and the PopTorch front end map a PyTorch model onto tiles, and the memory arithmetic that decides whether a model fits. It ends with failure modes, trade-offs and a checklist. Graphcore was acquired by SoftBank in July 2024, and the public PopTorch repository was last updated in October 2023, so read this as a guide to a real, deployed architecture and its ideas. It is not a buying guide for a current product line. For the other dataflow designs, compare Cerebras wafer-scale, the Groq LPU and the SambaNova RDU.

Why a processor without caches

A GPU keeps model weights and activations in HBM, a few gigabytes to a couple of hundred gigabytes of DRAM beside the die. Every kernel streams operands from HBM through L2 and L1 into registers. Bandwidth is high, but each byte travels a long way, and caches decide at run time what stays close. Most GPU optimisation work, such as fusion, tiling and FlashAttention, is about avoiding trips to HBM.

The IPU asks a different question. What if the working set lived entirely in SRAM next to the arithmetic? SRAM is fast and close but expensive in area, so it can only work if it is split into many small pieces, each owned by one core. That gives you a tile: a core plus 624 KiB of local memory. A tile can only load and store its own memory. If it needs data from another tile, the data must be sent explicitly over an on-chip fabric called the exchange.

Once each core sees only its own memory, there is nothing for a cache to do and nothing to keep coherent. Every movement is planned. This is why the IPU is called a graph processor: the program is a static graph of tensors, compute steps and copies, and the compiler places all three before the program runs.

The GC200 by the numbers

The figures below are for the Colossus MK2 GC200, the chip in the IPU-M2000 and IPU-POD systems. Throughput numbers are Graphcore's quoted peaks; use them only to compare orders of magnitude, not to predict a workload.

PropertyGC200What it means for software
Tiles1,472Tensors are split into up to 1,472 pieces
Worker threads6 per tile, 8,832 per chipEach tile interleaves 6 threads to hide pipeline latency
SRAM per tile624 KiBThe hard ceiling for one tile's code, data and temporaries
SRAM per chipabout 900 MBWeights, activations and optimizer state must fit, or be streamed
Quoted FP16 peak250 TFLOPS (Graphcore figure)Reached only by large, well-mapped matmuls
Chip-to-chipIPU-LinksUsed for model sharding, pipelining and replica collectives
Off-chip memoryStreaming memory (off-chip DRAM)Explicitly copied in and out, never cached

The Bow IPU that followed keeps the same 1,472-tile layout and adds a second wafer bonded to the logic wafer (TSMC wafer-on-wafer), mainly to improve power delivery. That lets it run at a higher clock, and Graphcore quoted 350 TFLOPS. The programming model did not change, so everything below applies to both chips.

Bulk synchronous execution

The IPU runs bulk synchronous parallel (BSP) programs. Each step has three phases. In the compute phase, every tile runs its assigned code on data already in its own memory, with no communication. In the sync phase, every tile waits at a hardware barrier until all tiles have finished. In the exchange phase, tiles send and receive the data they need for the next compute phase. The compiler fixed this schedule in advance.

BSP has three consequences that matter to anyone profiling an IPU program. First, a step is only as fast as its slowest tile, so an unevenly mapped tensor makes the whole chip wait. Second, exchange time is real time: if a layer needs a lot of data rearranged between tiles, the exchange phase can rival the compute phase. Third, because the schedule is static, there are no run-time races or cache misses to debug. If a program runs once, its timing is almost perfectly repeatable, which suits latency-sensitive inference.

GC200 IPU: 1,472 tiles, each a core with its own SRAM, joined by the exchangeTile624 KiBTile6 threadsTile624 KiBTile6 threadsTile624 KiBTile6 threadsTile624 KiBTile6 threadsTile624 KiBTile6 threadsTile624 KiBTile6 threads... 1,472 tilesabout 900 MB SRAM in totalno caches, no on-chip DRAMa tile reads only its own memoryIPU-Exchangeall-to-all tile fabric, schedule fixed at compile timeIPU-Linkschip-to-chip, for sharding and pipelinesHost CPUruns PopTorch, feeds dataStreaming memoryDRAM beside the chipsOther IPUsreplicas or pipeline stageshost linkexplicit copiesOne step = compute phase (tiles work alone) / sync (every tile waits) / exchange (data moves)computesyncexchangecomputesyncexchange
The tile array, the exchange fabric and the BSP cycle. Each tile owns its own SRAM; data reaches another tile only through a scheduled exchange phase.

Poplar: tensors, vertices and compute sets

Poplar is the C++ graph library underneath every IPU framework. A program has four parts. Tensors are variables that you map to tiles. Vertices are small C++ codelets that run on one tile. Compute sets are groups of vertices that run together in one compute phase. A program is a sequence of compute sets and copies. The example below follows Graphcore's own vertex tutorial. It puts one element of each tensor on each of four tiles, then runs a vertex per tile that sums a slice of the input.

// codelets.cpp: compiled for the tile processor, runs on one tile
#include <poplar/Vertex.hpp>
class SumVertex : public poplar::Vertex {
public:
  poplar::Input<poplar::Vector<float>> in;   // may live on other tiles; exchange brings it
  poplar::Output<float> out;                 // must live on this tile
  bool compute() {
    *out = 0;
    for (const auto &v : in) *out += v;
    return true;
  }
};

// host.cpp: builds and runs the graph
Graph graph(device.getTarget());
graph.addCodelets("codelets.cpp");
Tensor v1 = graph.addVariable(FLOAT, {4}, "v1");
Tensor v2 = graph.addVariable(FLOAT, {4}, "v2");
for (unsigned i = 0; i < 4; ++i) {
  graph.setTileMapping(v1[i], i);           // element i lives on tile i
  graph.setTileMapping(v2[i], i);
}
ComputeSet cs = graph.addComputeSet("sum");
for (unsigned i = 0; i < 4; ++i) {
  VertexRef vtx = graph.addVertex(cs, "SumVertex");
  graph.connect(vtx["in"], v1.slice(i, 4));  // reads tiles i..3: the compiler adds exchanges
  graph.connect(vtx["out"], v2[i]);
  graph.setTileMapping(vtx, i);
}
Sequence prog;
prog.add(Execute(cs));
Engine engine(graph, prog);
engine.load(device);
engine.run(0);

The important line is graph.connect(vtx["in"], v1.slice(i, 4)). The vertex on tile 0 reads elements that live on tiles 0 to 3, so the compiler must insert an exchange before the compute set runs. You never write that copy yourself, but you pay for it. In real models you rarely write vertices. The PopLibs libraries (poplin for matmuls, popnn for neural-network operations, popops for element-wise work) generate mapped tensors and vertex graphs for you.

PopTorch: training a PyTorch model

Most users reached the IPU through PopTorch, which traces a PyTorch module, lowers it to a Poplar graph and compiles it once. The training loss must be computed inside forward, because the whole step, forward, backward and optimizer, becomes one compiled program. Four options control how much data one host call consumes.

import torch
import poptorch

class Net(torch.nn.Module):
    def __init__(self):
        super().__init__()
        self.encoder = Encoder()          # your layers
        self.head = Head()
        self.loss = torch.nn.CrossEntropyLoss()

    def forward(self, x, labels=None):
        out = self.head(self.encoder(x))
        if labels is None:
            return out
        return out, self.loss(out, labels)   # loss inside forward: one compiled step

model = Net()
# Pipeline: encoder on IPU 0, head on IPU 1
model.encoder = poptorch.BeginBlock(model.encoder, ipu_id=0)
model.head = poptorch.BeginBlock(model.head, ipu_id=1)

opts = poptorch.Options()
opts.deviceIterations(16)                    # steps per host call
opts.Training.gradientAccumulation(8)        # micro-batches per weight update
opts.replicationFactor(4)                    # data-parallel copies of the 2-IPU pipeline
opts.setAvailableMemoryProportion({"IPU0": 0.3, "IPU1": 0.3})
opts.TensorLocations.setOptimizerLocation(
    poptorch.TensorLocationSettings().useOnChipStorage(False))  # Adam state off-chip

loader = poptorch.DataLoader(opts, dataset, batch_size=4, shuffle=True)
optimizer = poptorch.optim.AdamW(model.parameters(), lr=1e-4)
train = poptorch.trainingModel(model, options=opts, optimizer=optimizer)

for x, y in loader:
    out, loss = train(x, y)               # first call compiles; this can take minutes

Here is the batch arithmetic, which trips up almost everyone. The micro-batch is 4 samples. Gradient accumulation of 8 means one weight update sees 4 × 8 = 32 samples per replica. Four replicas make the global batch 128. Device iterations of 16 means a single host call runs 16 such steps, so the loader yields 4 × 8 × 4 × 16 = 2,048 samples per call. If you tune the learning rate for a global batch of 128 but read the loader size as the batch, you will mis-set the schedule by a factor of 16. Pipelined training also requires gradient accumulation to be at least the number of pipeline stages, counting forward and backward stages, or the pipeline never fills.

Worked example: will it fit?

Because there is no HBM, the first design question for any model is whether its working set fits in roughly 900 MB per chip, and every byte counts. Work through a BERT-Large-sized encoder with 340 million parameters, trained with AdamW in mixed precision.

ItemBytes per parameterTotalFits in one IPU?
FP16 weights2680 MBBarely, leaving no room for anything else
FP16 gradients2680 MBNo
FP32 master weights + Adam m and v124.08 GBNo

The answer is a combination of three techniques. Pipelining across four IPUs puts a quarter of the layers on each chip, so each holds about 170 MB of FP16 weights. Off-chip optimizer state keeps the 4 GB of Adam state in streaming memory and moves it in only for the weight update, which runs once per accumulated batch, so its transfer cost is spread over many micro-batches. Recomputation stores only checkpoint activations between stages and recomputes the rest in the backward pass. What is left is code, exchange buffers and temporaries. Code is not free on an IPU: each tile stores the vertex code it runs, and a model with many different operations can spend a large share of its SRAM on instructions.

setAvailableMemoryProportion is the main control knob. It caps the fraction of tile memory that matmul and convolution planners may use for temporaries. Lower it and the planner picks slower plans that use less memory. Raise it and you gain speed until the graph no longer fits and compilation fails. Graphcore's PopVision Graph Analyser shows memory per tile after compilation. Use it before guessing.

Failure modes

These are the failures teams actually met on IPUs, roughly in order of frequency.

  • Out of memory at compile time. A tile ran out of its 624 KiB, often just one tile holding an unevenly mapped embedding or the vertex code for a rare operation. Read the per-tile memory report, not the chip total. Reduce available memory proportion, recompute more, or move a layer to another pipeline stage.
  • Recompilation on every shape change. The graph is static, so a new sequence length or batch size is a new program. Pad or bucket inputs to a few fixed shapes and cache the compiled executables. Never let a variable-length tail batch reach the model.
  • Pipeline imbalance. One stage takes twice as long as the others, and BSP makes every stage wait for it. Profile per-stage cycle counts and move layers between IPUs until the stages are within about 10% of each other.
  • Host-bound training. Too few device iterations means the chip waits on Python and the data loader between calls. Raise device iterations and use the asynchronous data loader mode.
  • Unsupported operations. Dynamic control flow, data-dependent shapes and some sparse operations either fail to lower or fall back to slow forms. Find them with a small-scale compile before committing to a port.
  • A stale software stack. PopTorch is pinned to an older PyTorch release, so newer model code may need operators it lacks. Pin the SDK, PyTorch and Python together.

Trade-offs

ChooseWhenCost
IPU, model fits in SRAMSmall or medium models, irregular or fine-grained compute, graph and sparse workloads, latency-critical inference with fixed shapesStatic shapes, long compiles, a narrower software ecosystem
IPU, pipelined across chipsModels of a few hundred million to a few billion parameters with a stable architectureManual stage balancing, gradient-accumulation constraints
GPULarge language models, rapidly changing model code, dynamic shapes, anything that depends on the CUDA library ecosystemMemory-bound kernels, cache behaviour that is harder to predict
TPU or other dataflow chipsVery large dense training with compiler-friendly models; see the TPU guideA different compiler and its own static-shape rules

The lesson: compile-time data movement beats caches when movement is predictable, and costs you wherever it is not.

What to do next

  1. If you have no hardware, read the Poplar tutorials in graphcore/examples (tutorials were moved there from the archived graphcore/tutorials repo) and run them on the IPU Model, the CPU-based simulator, to learn tile mapping and exchanges.
  2. For any model you plan to port, write its memory table first, covering weights, gradients, optimizer state and activations per pipeline stage, and compare it with about 900 MB per chip.
  3. Compile a one-layer version before porting the whole model, and list every operation that fails to lower or falls back.
  4. Fix input shapes. Choose a few sequence-length buckets and confirm each compiles once and is cached.
  5. Write down the four batch numbers (micro-batch, gradient accumulation, replicas, device iterations) and derive the global batch and the loader batch before tuning any learning rate.
  6. Open the compiled report in PopVision, check the busiest tile's memory and the per-stage cycle balance, and adjust available memory proportion or stage boundaries.
  7. Pin the SDK, PyTorch and Python versions in a container image, because the toolchain is no longer moving with upstream PyTorch.
Key takeaway: The IPU splits one chip into 1,472 tiles, each owning 624 KiB of SRAM. It runs compute, sync and exchange phases that the compiler schedules ahead of time. Anything that fits in about 900 MB per chip runs with predictable timing. Larger models need pipelining, off-chip optimizer state and recomputation, and every change of shape means a recompile. Budget memory per tile, fix shapes and derive the real batch size before tuning.