Every accelerator design is an answer to one question: where do the bytes live, and how far do they travel before they meet an arithmetic unit? A GPU keeps model weights in HBM stacks beside the die and pulls them through a few terabytes per second of memory bandwidth. Cerebras took the opposite extreme. Instead of cutting a silicon wafer into hundreds of chips, it keeps the wafer whole and turns it into one processor whose memory is static RAM spread across the entire surface, next to the cores that use it.
The third generation, the Wafer-Scale Engine 3 (WSE-3), sits inside the CS-3 system. This article explains what is physically on the wafer, why that layout changes which bottleneck you hit, how training and inference are actually mapped onto it, how software reaches it at three levels of abstraction, and where it fails or costs you more than a GPU cluster. Figures quoted are Cerebras's published specifications; anything Cerebras has not published is marked as an estimate.
One wafer, one processor
A standard 300 mm wafer normally yields dozens of large GPU dies. The WSE-3 uses the largest square that fits on the wafer as a single device. Cerebras lists these headline numbers for it:
| Property | WSE-3 (published) | What it means for software |
|---|---|---|
| Silicon area | 46,225 mm2, TSMC 5 nm | Roughly 57 times the area of a large GPU die |
| Transistors | About 4 trillion | Most of them are SRAM and routers, not math |
| Cores | 900,000 processing elements | Each runs its own small program |
| On-chip memory | 44 GB SRAM | About 48 KB per core, one cycle away |
| Memory bandwidth | 21 PB/s | Aggregate of every core reading its own SRAM |
| Fabric bandwidth | 214 Pb/s | Core-to-core mesh links summed |
| Peak AI compute | 125 PF | Vendor peak; real kernels reach a fraction |
Two engineering problems had to be solved to make this work. The first is defects: every wafer has some, and a chip the size of a wafer will always contain several. Cerebras builds in spare cores and spare links and routes around faulty ones at configuration time, so the software sees a clean rectangular mesh even though the physical silicon is not perfect. The second is that a wafer cannot be packaged like a normal chip; the CS-3 is a custom chassis with its own power delivery and cooling. That is why you buy or rent a whole system, never a card.
Inside a processing element
The useful mental model is a two-dimensional grid of tiny computers. Each processing element (PE) has a SIMD-capable arithmetic unit, about 48 KB of SRAM that holds both its code and its data, and a router connected to its four neighbours. There is no shared cache and no global memory: a PE can only read its own SRAM. Anything else arrives as a message over the mesh.
Those messages are called wavelets. A program configures circuits through the mesh, and each circuit is identified by a colour, a virtual channel bound to routing resources. Each router supports a limited number of colours (the SDK documentation describes 24 routable ones per PE), so laying out a computation means deciding which data flows on which colour across which rows and columns. Arrival of a wavelet can trigger a task on the receiving PE. That is what dataflow means here: work is scheduled by data turning up, not by a central instruction stream.
Compare that with a GPU streaming multiprocessor, which hides memory latency by keeping dozens of warps in flight and switching between them. A PE has nothing to hide: its operand is in local SRAM. The cost moves from latency to placement. If two operands that must meet sit on opposite sides of the wafer, the mesh carries them hop by hop, and the compiler's main job is keeping that distance short. Cerebras also says the dataflow design can skip multiplications by zero, which lets unstructured sparsity save work, something dense GPU tensor cores cannot exploit in general.
Why SRAM changes the bottleneck
Why go to this trouble? Because large-model inference is mostly a memory problem. Generating one token in a dense transformer reads every weight once. On a single H100 SXM with 3.35 TB/s of HBM3 bandwidth, a model whose weights total 140 GB would take at least 42 ms per token just to stream them, before any arithmetic or communication, and it does not even fit in 80 GB. Batching amortises those reads across users, which is why GPU serving is a throughput game and per-user speed is capped by bandwidth.
The WSE-3 moves the weights into SRAM next to the cores. Its 21 PB/s is more than 6,000 times one H100's HBM bandwidth. The catch is capacity: 44 GB holds about 22 billion parameters at 16-bit precision, fewer once activations and a KV cache claim their share. So the architecture forces a choice that shapes everything above it: either keep the model on the wafer by spreading it across several wafers (inference), or keep only the working set on the wafer and stream the weights in from outside (training).
Training: weight streaming with MemoryX and SwarmX
For training, Cerebras uses an execution mode it calls weight streaming. Model weights and optimizer state live in an external memory service, MemoryX. Activations live on the wafer. Training proceeds one layer at a time: the layer's weights are streamed onto the wafer, the forward pass for the whole batch runs, and later the backward pass streams the weights in again and streams gradients out. The optimizer update runs in MemoryX, not on the wafer. A second service, SwarmX, broadcasts weights to several CS-3 systems and sums their gradients on the way back, which is plain data parallelism.
The consequence for a training engineer is that model size and compute size are decoupled. A bigger model needs more MemoryX capacity, not a different parallelism plan; more throughput needs more CS-3 systems, still in plain data parallelism. There is no tensor parallelism and no pipeline schedule to tune, which on GPU clusters is a large share of the engineering effort. In pseudocode one step looks like this:
# One weight-streaming training step (conceptual, not a Cerebras API)
for layer in model.layers: # forward
w = memoryx.read(layer.id) # streamed onto every wafer via SwarmX
acts[layer.id] = wafer.forward(layer, w, acts[layer.id - 1]) # activations stay on wafer
loss = wafer.loss(acts[-1], labels)
for layer in reversed(model.layers): # backward
w = memoryx.read(layer.id) # weights stream in again
grad_w, grad_in = wafer.backward(layer, w, acts[layer.id], grad_in)
swarmx.reduce_to(memoryx, layer.id, grad_w) # sum across CS-3s, then send out
memoryx.optimizer_step() # Adam etc. runs off-waferThe constraint is that every layer's weights cross the link into the wafer twice per step. That traffic is amortised across the batch, so weight streaming wants large batches and long sequences, just as GPU data parallelism does. Small-batch fine-tuning of a huge model is where this mode is least efficient.
Inference: spreading a model across wafers
Inference uses the other side of the trade. To get the bandwidth win, the weights must already be on silicon, so a large model is partitioned by layers across several wafers and tokens flow through them as a pipeline, the same idea as pipeline parallelism on GPUs but with each stage holding its weights in SRAM. Cerebras has not published the exact mapping it uses for each hosted model, so the worked example below is arithmetic, not a description of its deployment.
Take a 70-billion-parameter dense model at 16-bit weights: 140 GB. At 44 GB per wafer, the weights alone need at least four wafers, and in practice more, because the KV cache, activations and the code in each PE also occupy SRAM. Suppose five wafers each hold 28 GB of weights. Each wafer reads its slice in 28 GB / 21 PB/s, about 1.3 microseconds. Five stages give a memory-time floor near 7 microseconds per token, which would allow over 100,000 tokens per second for one stream. No system reaches that, and that is the lesson: once weights sit in SRAM, the binding constraints become arithmetic, hops across the mesh, and the latency of passing activations between wafers. Those are what set the real per-user speed, which Cerebras's own benchmarks place far above GPU servers for a single stream.
The flip side is the KV cache. Long contexts and many concurrent users need cache memory that competes with weights for the same fixed SRAM. A GPU server can trade batch size against HBM headroom freely; a wafer pipeline has much less slack, so the number of concurrent long-context requests per wafer is the figure to ask a vendor about.
Three ways to program it
There are three ways to reach a WSE-3, and they demand very different effort.
- Hosted inference. Cerebras runs an API that follows the OpenAI chat completions shape, so existing clients work with a different base URL. This is how most teams meet the hardware.
- Model training on a CS-3 cluster. You write PyTorch models in the Cerebras software stack, which compiles the graph for the wafer and runs it in weight streaming mode. Supported layer types and shapes are what the compiler knows; a custom CUDA kernel has no equivalent.
- The SDK and CSL. For HPC and research kernels you program the PEs directly in the Cerebras Software Language, a C-like language based on Zig, with a host program in Python that copies data in and launches functions.
A hosted call looks like any OpenAI-style request:
import os
from openai import OpenAI
client = OpenAI(base_url="https://api.cerebras.ai/v1",
api_key=os.environ["CEREBRAS_API_KEY"])
MODEL = os.environ["CEREBRAS_MODEL"] # pick a current ID from the provider's model list
resp = client.chat.completions.create(
model=MODEL,
messages=[{"role": "user", "content": "Explain weight streaming in two sentences."}],
max_tokens=200,
)
print(resp.choices[0].message.content)
print(resp.usage) # log tokens so you can compute cost per requestAt the other end, a minimal CSL program declares the rectangle of PEs it uses, assigns code to each tile, and exports symbols the host can reach. The layout file imports the <memcpy/get_params> module so the host can move data:
// layout.csl: one PE, one exported array, one exported function
const memcpy = @import_module("<memcpy/get_params>", .{ .width = 1, .height = 1 });
layout {
@set_rectangle(1, 1);
@set_tile_code(0, 0, "pe_program.csl", .{ .memcpy_params = memcpy.get_params(0) });
@export_name("A", [*]f32, true);
@export_name("scale", fn()void);
}
// pe_program.csl: runs on that PE
param memcpy_params: comptime_struct;
const sys_mod = @import_module("<memcpy/memcpy>", memcpy_params);
var A = @zeros([4]f32);
var ptr_A: [*]f32 = &A;
fn scale() void {
for (@range(i16, 4)) |i| { A[i] = A[i] * 2.0; }
sys_mod.unblock_cmd_stream(); // tell the host the launched function finished
}
comptime {
@export_symbol(ptr_A, "A");
@export_symbol(scale);
}The host side uses SdkRuntime to load the compiled program, memcpy_h2d to fill A, launch to call scale and memcpy_d2h to read the result back. Real kernels repeat that pattern over thousands of PEs, with colours carrying data between them. Most teams never write CSL; it matters when your workload is a stencil or sparse solver that no framework maps well.
Failure modes
The ways a wafer-scale project goes wrong are mostly about fit:
- The model does not fit the compiler. Dynamic shapes, unusual attention variants or custom ops can fail to compile or fall back to slow paths. Test your exact architecture before committing, not a close cousin.
- Small batches in training. Weight streaming moves every layer twice per step; tiny batches leave the wafer waiting on MemoryX.
- Context length surprises in inference. KV cache competes with weights for fixed SRAM, so long-context concurrency can be lower than short-prompt benchmarks suggest. Benchmark at your real prompt and output lengths.
- Numerics differences. Different reduction orders and formats mean results differ slightly from a GPU run. Compare evaluation metrics, not bitwise outputs.
- Lock-in. Training code written for the Cerebras stack does not run elsewhere unchanged. Keep model definitions and data pipelines portable and checkpoints in a standard format.
- Capacity. Few organisations own CS-3 systems; most access is through Cerebras's own cloud or partners, so rate limits and regional availability belong in your risk list.
Trade-offs
| Dimension | WSE-3 / CS-3 | GPU cluster (H100 class) |
|---|---|---|
| Per-stream decode speed | Very high; weights in SRAM | Bounded by HBM bandwidth per token |
| Memory capacity per device | 44 GB SRAM | 80 GB HBM3 per H100, much more per node |
| Training parallelism | Data parallel only, via weight streaming | Data, tensor, pipeline, expert; tuned by hand |
| Programmability | Compiler-supported graphs; CSL for custom | CUDA, Triton, any kernel you write |
| Ecosystem | One vendor | Every framework, every cloud |
| Unit of purchase | Whole system or hosted API | Single GPU hour upward |
The short version: the wafer wins when single-stream latency or simple scaling of a supported model is worth more than flexibility. For a different take on the same SRAM-first idea, compare the Groq LPU, which uses many smaller chips and a static schedule, and the SambaNova RDU, which keeps HBM and DDR tiers. The TPU and the H100 are the baselines most comparisons start from.
What to do next
- Write down your bottleneck first: per-user latency, aggregate throughput or training time. Wafer-scale helps the first one most.
- Run your real prompts through the hosted API with the OpenAI client above, logging time to first token, tokens per second and token usage for each request.
- Benchmark the same prompts on your current GPU serving stack at equal output length, and compare cost per million tokens and p95 latency, not peak numbers.
- If you are evaluating training, confirm that your exact architecture compiles and ask for measured throughput at your batch size and sequence length.
- Keep model code and checkpoints portable so you can move back to GPUs without a rewrite.
- Only consider CSL if your kernel is a stencil, sparse solver or similar workload that frameworks map poorly, and start in the SDK simulator.