Groq's Language Processing Unit (LPU) is an inference accelerator built on a different bet from a GPU. A GPU keeps model weights in off-chip HBM and hides latency with thousands of threads, caches and a hardware scheduler that decides at run time what executes next. The LPU keeps weights in on-chip SRAM, has no caches and no dynamic scheduler, and moves the whole scheduling problem into the compiler, which decides on which clock cycle every instruction runs and every operand arrives. The result is hardware whose latency is deterministic and whose per-chip memory bandwidth is far higher than HBM, paid for with very little memory per chip.
This article explains the design in terms of what software sees, and how to decide whether an LPU-backed endpoint fits your workload. Specifications are the ones Groq published for its first-generation GroqChip and the ISCA 2020 and 2022 papers on the Tensor Streaming Processor (TSP) the LPU is built on. Newer generations change the numbers, so treat them as a worked baseline, not a current datasheet.
Why decode rewards memory bandwidth
Generating one token with a dense transformer means multiplying the current activation vector by every weight matrix in the model. At batch size one that is a matrix-vector product: each weight byte is read once and used for roughly two floating-point operations. The arithmetic intensity is about 2 FLOPs per byte, far below what any modern accelerator needs to be compute-bound, so decode speed is set by how fast weights can be streamed out of memory. The decode math article works this roofline argument through for GPUs.
A rough lower bound on time per token is therefore model bytes divided by memory bandwidth. An H100 SXM has about 3.35 TB/s of HBM3 bandwidth, so a 70 GB set of weights (70 billion parameters at one byte each) needs at least about 21 ms per token on one GPU, or about 2.6 ms if eight GPUs split the weights with tensor parallelism and read in parallel. GPUs recover efficiency by batching: one weight read serves every sequence in the batch, so throughput rises while per-token latency stays near the bandwidth floor. The HBM article covers why off-chip bandwidth is hard to raise.
The LPU attacks the floor itself. On-die SRAM sits next to the compute units, so the first-generation GroqChip quotes about 80 TB/s of on-die memory bandwidth, roughly 24 times one H100's HBM. The catch is capacity: about 230 MB of SRAM per chip, against 80 GB of HBM on an H100. A 70B model does not fit on one chip or ten; it needs hundreds. Every other design decision follows from that trade.
Inside the chip: the tensor streaming processor
A GPU tiles the die with many identical cores, each with its own mix of ALUs, registers and cache. The TSP turns that inside out. Units of one kind are grouped into a slice: matrix units (MXM), switch and permute units (SXM), memory slices (MEM) and vector units (VXM). The slices sit side by side across the chip, and data flows horizontally between them as streams, advancing one slice per clock cycle in either direction. An instruction tells a slice to read a stream, operate on it and write the result to an outgoing stream at a particular cycle.
The published first-generation figures are 750 TOPS at INT8 and 188 TFLOPS at FP16, from a 320 by 320 fused dot-product matrix unit and 5,120 vector ALUs. Vectors are 320 elements wide, organised as 20 superlanes of 16 lanes.
The most important feature for a software engineer is what is missing. There is no cache hierarchy, so no hit or miss and no eviction. There is no hardware arbitration between competing requests, so no contention-dependent latency. There is no out-of-order execution and no warp scheduler. Each of those mechanisms exists on a GPU to cope with uncertainty at run time. The TSP removes the uncertainty instead, which is only possible because a neural network's dataflow graph is known in advance.
The compiler is the scheduler
Because the hardware never decides anything at run time, the compiler must decide everything at compile time: which memory slice holds each weight tile, which cycle each stream is read, how vectors are routed through the switch units, and when results reach the next chip. The output is closer to a cycle-by-cycle timetable than to a GPU kernel launch sequence. Running the same program twice takes the same number of cycles, so the latency of a forward pass is a property of the compiled program rather than a distribution measured under load.
That has direct consequences for the software that sits on top:
- Latency is predictable per step. Tail latency inside a forward pass does not grow with contention because there is none. What still varies is queueing in front of the system, prompt length and output length.
- Shapes must be known. A static schedule is compiled for particular tensor shapes. Variable lengths generally mean compiling for fixed sizes and padding or bucketing, the same compromise XLA makes for TPUs (see the TPU article).
- No user kernels. On a GPU you can drop into CUDA or Triton to fix a slow operator. On the LPU the compiler owns placement and timing, so new model architectures arrive when the toolchain supports them, and hosted customers choose from the models the provider has compiled.
Scaling out: hundreds of chips as one pipeline
Since one chip holds about 230 MB, a large model is partitioned across many chips, and each chip keeps its share of the weights resident in SRAM for the life of the deployment. Only activations move between chips. The TSP scale-out design extends static scheduling across chip boundaries: the 2022 ISCA paper describes chip-to-chip links that are scheduled by software, with clocks kept aligned closely enough (Groq calls this plesiosynchronous) that the compiler can predict when a vector sent from one chip arrives at the next.
The simplest mental model is a pipeline, possibly with several chips sharing each stage. Pipeline parallelism has a well-known cost: a single token must pass through every stage in turn, so at batch size one most chips are idle most of the time. The fix is the same as on GPUs: keep many sequences in flight so each stage works on a different token. The difference is that the per-stage time is tiny and fixed, so the pipeline can be very deep without the jitter that makes deep GPU pipelines painful.
The KV cache is the second capacity pressure. Every in-flight sequence keeps keys and values for every layer and every token of context, and on an SRAM-only machine that state competes with weights for the same scarce bytes. Long contexts and large batches are therefore more expensive on an LPU than on an HBM machine, where the cache is the main thing the 80 GB is for.
Worked example: sizing a 70B model
Size a deployment of a Llama-2-70B-class model with grouped-query attention: 80 layers, 8 key-value heads of dimension 128, weights quantised to one byte per parameter, KV cache in 16-bit. The script computes the chip count for weights, the KV cache cost per sequence, and a memory-traffic estimate per token for the LPU layout and for H100s. These count memory traffic only; real systems add interconnect hops, activations, compute and software overhead.
import math
PARAMS = 70e9 # dense parameters
W_BYTES = 1 # INT8 / FP8 weights
LAYERS = 80
KV_HEADS = 8
HEAD_DIM = 128
KV_BYTES = 2 # FP16 cache
SRAM_PER_CHIP = 230e6 # first-generation GroqChip
SRAM_BW = 80e12 # bytes/s, on-die
HBM_BW = 3.35e12 # bytes/s, one H100 SXM
weights = PARAMS * W_BYTES
chips_for_weights = math.ceil(weights / SRAM_PER_CHIP)
kv_per_token = 2 * LAYERS * KV_HEADS * HEAD_DIM * KV_BYTES # K and V
kv_4k = kv_per_token * 4096
# Pipeline: a token streams each chip's slice in turn, so the chip times add up.
lpu_pipeline = chips_for_weights * (SRAM_PER_CHIP / SRAM_BW)
gpu1_floor = weights / HBM_BW
gpu8_floor = weights / (8 * HBM_BW)
print(f"chips for weights only : {chips_for_weights}")
print(f"KV per token : {kv_per_token/1024:.0f} KiB")
print(f"KV per 4k sequence : {kv_4k/1e9:.2f} GB = {kv_4k/SRAM_PER_CHIP:.1f} chips of SRAM")
print(f"LPU pipeline estimate : {lpu_pipeline*1e3:.2f} ms")
print(f"1x H100 floor : {gpu1_floor*1e3:.1f} ms")
print(f"8x H100 TP floor : {gpu8_floor*1e3:.2f} ms")The output is 305 chips for the weights alone, 320 KiB of KV cache per token, 1.34 GB (about 5.8 chips of SRAM) for one 4,096-token sequence, a pipeline-only LPU estimate of about 0.88 ms per token, 20.9 ms on one H100 and 2.61 ms on eight. Three lessons fall out of it. First, the single-stream speed advantage is real and comes straight from bandwidth: the estimate is about three times lower than an eight-GPU node, and splitting each stage across chips in parallel lowers it further at the cost of more chip-to-chip traffic. Second, the hardware footprint is large: hundreds of chips across several racks for one model replica. Third, every extra concurrent long-context sequence costs several chips' worth of SRAM, so concurrency and context length, not weights, often decide the real chip count. Change the parameters to your model and context before you believe any vendor comparison.
Using an LPU from software
Most teams meet the LPU through GroqCloud, Groq's hosted API, which is compatible with the OpenAI chat completions format. Groq also publishes a Python SDK. Either works; the OpenAI-compatible route lets you switch an existing client by changing the base URL. Model IDs change as models are added and retired, so read them from the models endpoint rather than hard-coding a name from a blog post.
import os, time
from openai import OpenAI
client = OpenAI(api_key=os.environ["GROQ_API_KEY"],
base_url="https://api.groq.com/openai/v1")
model = os.environ["GROQ_MODEL"] # pick one from client.models.list()
t0 = time.perf_counter()
first = None
chunks = 0
stream = client.chat.completions.create(
model=model,
messages=[{"role": "user", "content": "Explain KV caching in 3 sentences."}],
max_tokens=200,
stream=True,
)
for event in stream:
if event.choices and event.choices[0].delta.content:
first = first or time.perf_counter()
chunks += 1
end = time.perf_counter()
print(f"time to first token {first - t0:.3f}s, "
f"{chunks / (end - first):.0f} chunks/s after first token")Measure two numbers separately: time to first token, which includes network, queueing and prompt processing, and inter-token rate after the first token, which reflects the decode path the hardware is optimised for. Streaming chunks are not exactly tokens, so use a non-streaming call and its usage field for exact token counts. Run the same harness against your current provider at the same prompt and output lengths before drawing conclusions; the prefill and decode split article explains why the two phases behave differently.
On the hardware roadmap: in December 2025 a non-exclusive licensing agreement giving Nvidia access to Groq's inference technology was announced, members of Groq's leadership joined Nvidia, and Groq continues to run GroqCloud. Nvidia has since announced LPX racks of 256 Groq 3 LPUs, positioned to handle decode alongside its GPUs. Press coverage of the per-chip specifications is inconsistent, so take figures for newer generations from Nvidia's or Groq's own published specifications rather than from summaries.
Failure modes
- Assuming fast decode means fast end-to-end. Long prompts are dominated by prefill, and queueing in front of a shared service adds latency the chip never sees. Measure time to first token on your real prompt lengths.
- Quantisation drift. Fitting weights into SRAM pushes providers toward low-precision formats. Run your own evaluation set against the hosted model rather than trusting the original model's benchmark scores.
- Context and concurrency limits. SRAM-resident KV cache makes very long contexts expensive. Check the context window and rate limits the endpoint actually offers for each model.
- Misreading determinism. Deterministic timing is not deterministic output; sampling above temperature zero still varies.
Trade-offs
| Dimension | LPU (SRAM, static schedule) | GPU (HBM, dynamic schedule) |
|---|---|---|
| Per-stream decode speed | Very high: bandwidth floor far below HBM | Bounded by HBM bandwidth per GPU group |
| Memory per chip | About 230 MB (first generation) | 80 GB or more |
| Chips per large model | Hundreds | One to eight per replica |
| Latency variance | Fixed per compiled step | Varies with load and contention |
| Flexibility | Compiler-owned; fixed shapes; no custom kernels | Any model, custom kernels, dynamic shapes |
| Long context, big batches | Costly: KV competes with weights for SRAM | Natural fit: HBM holds large caches |
| Training | Not the target workload | Primary workload |
An LPU fits interactive decode on popular models, such as voice agents and agent loops where each step waits on the last. A GPU is safer for custom models, very long contexts, training or offline throughput. Speculative decoding is the GPU-side answer to the same latency problem; benchmark it too.
What to do next
- Run the sizing script with your model's layer count, KV heads, precision and target context, and note whether weights or KV cache dominate the footprint.
- Build the streaming benchmark against GroqCloud and your current provider, and record time to first token and inter-token rate at your real prompt and output lengths.
- Run your own evaluation set on the hosted model to catch quality changes from quantisation before you switch traffic.
- List the steps in your product where a user or an agent waits on a sequential LLM call; those are where per-stream decode speed pays off.
- Keep a GPU fallback path for models and context lengths the LPU endpoint does not offer, and route by model and length.
- Recheck newer-generation specifications from Groq's and Nvidia's own published documents before any capacity plan.