Groq's Language Processing Unit (LPU) is an inference accelerator built on a different bet from a GPU. A GPU keeps model weights in off-chip HBM and hides latency with thousands of threads, caches and a hardware scheduler that decides at run time what executes next. The LPU keeps weights in on-chip SRAM, has no caches and no dynamic scheduler, and moves the whole scheduling problem into the compiler, which decides on which clock cycle every instruction runs and every operand arrives. The result is hardware whose latency is deterministic and whose per-chip memory bandwidth is far higher than HBM, paid for with very little memory per chip.

This article explains the design in terms of what software sees, and how to decide whether an LPU-backed endpoint fits your workload. Specifications are the ones Groq published for its first-generation GroqChip and the ISCA 2020 and 2022 papers on the Tensor Streaming Processor (TSP) the LPU is built on. Newer generations change the numbers, so treat them as a worked baseline, not a current datasheet.

Why decode rewards memory bandwidth

Generating one token with a dense transformer means multiplying the current activation vector by every weight matrix in the model. At batch size one that is a matrix-vector product: each weight byte is read once and used for roughly two floating-point operations. The arithmetic intensity is about 2 FLOPs per byte, far below what any modern accelerator needs to be compute-bound, so decode speed is set by how fast weights can be streamed out of memory. The decode math article works this roofline argument through for GPUs.

A rough lower bound on time per token is therefore model bytes divided by memory bandwidth. An H100 SXM has about 3.35 TB/s of HBM3 bandwidth, so a 70 GB set of weights (70 billion parameters at one byte each) needs at least about 21 ms per token on one GPU, or about 2.6 ms if eight GPUs split the weights with tensor parallelism and read in parallel. GPUs recover efficiency by batching: one weight read serves every sequence in the batch, so throughput rises while per-token latency stays near the bandwidth floor. The HBM article covers why off-chip bandwidth is hard to raise.

The LPU attacks the floor itself. On-die SRAM sits next to the compute units, so the first-generation GroqChip quotes about 80 TB/s of on-die memory bandwidth, roughly 24 times one H100's HBM. The catch is capacity: about 230 MB of SRAM per chip, against 80 GB of HBM on an H100. A 70B model does not fit on one chip or ten; it needs hundreds. Every other design decision follows from that trade.

Inside the chip: the tensor streaming processor

One GroqChip: functional slices side by side, data streams flowing across themMXMmatrixSXMswitchMEMSRAM slicesVXMvector ALUsMEMSRAM slicesSXMswitchMXMmatrixstreams flow east and west, one hop per cycleInstruction controlper-slice queues, issued on compiler-chosen cyclesNo caches, no arbitration, no out-of-order issue: every operand arrival time is known when the program is compiled.
The TSP layout: functional units are grouped into vertical slices by type, and data moves horizontally between them as streams.

A GPU tiles the die with many identical cores, each with its own mix of ALUs, registers and cache. The TSP turns that inside out. Units of one kind are grouped into a slice: matrix units (MXM), switch and permute units (SXM), memory slices (MEM) and vector units (VXM). The slices sit side by side across the chip, and data flows horizontally between them as streams, advancing one slice per clock cycle in either direction. An instruction tells a slice to read a stream, operate on it and write the result to an outgoing stream at a particular cycle.

The published first-generation figures are 750 TOPS at INT8 and 188 TFLOPS at FP16, from a 320 by 320 fused dot-product matrix unit and 5,120 vector ALUs. Vectors are 320 elements wide, organised as 20 superlanes of 16 lanes.

The most important feature for a software engineer is what is missing. There is no cache hierarchy, so no hit or miss and no eviction. There is no hardware arbitration between competing requests, so no contention-dependent latency. There is no out-of-order execution and no warp scheduler. Each of those mechanisms exists on a GPU to cope with uncertainty at run time. The TSP removes the uncertainty instead, which is only possible because a neural network's dataflow graph is known in advance.

The compiler is the scheduler

Because the hardware never decides anything at run time, the compiler must decide everything at compile time: which memory slice holds each weight tile, which cycle each stream is read, how vectors are routed through the switch units, and when results reach the next chip. The output is closer to a cycle-by-cycle timetable than to a GPU kernel launch sequence. Running the same program twice takes the same number of cycles, so the latency of a forward pass is a property of the compiled program rather than a distribution measured under load.

That has direct consequences for the software that sits on top:

  • Latency is predictable per step. Tail latency inside a forward pass does not grow with contention because there is none. What still varies is queueing in front of the system, prompt length and output length.
  • Shapes must be known. A static schedule is compiled for particular tensor shapes. Variable lengths generally mean compiling for fixed sizes and padding or bucketing, the same compromise XLA makes for TPUs (see the TPU article).
  • No user kernels. On a GPU you can drop into CUDA or Triton to fix a slow operator. On the LPU the compiler owns placement and timing, so new model architectures arrive when the toolchain supports them, and hosted customers choose from the models the provider has compiled.

Scaling out: hundreds of chips as one pipeline

A 70B model laid out as a pipeline of chips: weights stay put, activations moveChips 1-38embed + layers 1-10Chips 39-76layers 11-20...layers 21-70Last chipslayers 71-80 + LM headnext tokenScheduled chip-to-chip linkstransfer cycles fixed at compile timeEach chip holds one slice of the weights in SRAM. A token visits every chip once,so the system keeps many sequences in flight to keep the whole pipeline busy.
Pipeline layout for a large model across many chips. The number of chips per stage is illustrative.

Since one chip holds about 230 MB, a large model is partitioned across many chips, and each chip keeps its share of the weights resident in SRAM for the life of the deployment. Only activations move between chips. The TSP scale-out design extends static scheduling across chip boundaries: the 2022 ISCA paper describes chip-to-chip links that are scheduled by software, with clocks kept aligned closely enough (Groq calls this plesiosynchronous) that the compiler can predict when a vector sent from one chip arrives at the next.

The simplest mental model is a pipeline, possibly with several chips sharing each stage. Pipeline parallelism has a well-known cost: a single token must pass through every stage in turn, so at batch size one most chips are idle most of the time. The fix is the same as on GPUs: keep many sequences in flight so each stage works on a different token. The difference is that the per-stage time is tiny and fixed, so the pipeline can be very deep without the jitter that makes deep GPU pipelines painful.

The KV cache is the second capacity pressure. Every in-flight sequence keeps keys and values for every layer and every token of context, and on an SRAM-only machine that state competes with weights for the same scarce bytes. Long contexts and large batches are therefore more expensive on an LPU than on an HBM machine, where the cache is the main thing the 80 GB is for.

Worked example: sizing a 70B model

Size a deployment of a Llama-2-70B-class model with grouped-query attention: 80 layers, 8 key-value heads of dimension 128, weights quantised to one byte per parameter, KV cache in 16-bit. The script computes the chip count for weights, the KV cache cost per sequence, and a memory-traffic estimate per token for the LPU layout and for H100s. These count memory traffic only; real systems add interconnect hops, activations, compute and software overhead.

import math

PARAMS        = 70e9      # dense parameters
W_BYTES       = 1         # INT8 / FP8 weights
LAYERS        = 80
KV_HEADS      = 8
HEAD_DIM      = 128
KV_BYTES      = 2         # FP16 cache

SRAM_PER_CHIP = 230e6     # first-generation GroqChip
SRAM_BW       = 80e12     # bytes/s, on-die
HBM_BW        = 3.35e12   # bytes/s, one H100 SXM

weights = PARAMS * W_BYTES
chips_for_weights = math.ceil(weights / SRAM_PER_CHIP)

kv_per_token = 2 * LAYERS * KV_HEADS * HEAD_DIM * KV_BYTES   # K and V
kv_4k = kv_per_token * 4096

# Pipeline: a token streams each chip's slice in turn, so the chip times add up.
lpu_pipeline = chips_for_weights * (SRAM_PER_CHIP / SRAM_BW)
gpu1_floor = weights / HBM_BW
gpu8_floor = weights / (8 * HBM_BW)

print(f"chips for weights only : {chips_for_weights}")
print(f"KV per token           : {kv_per_token/1024:.0f} KiB")
print(f"KV per 4k sequence     : {kv_4k/1e9:.2f} GB = {kv_4k/SRAM_PER_CHIP:.1f} chips of SRAM")
print(f"LPU pipeline estimate  : {lpu_pipeline*1e3:.2f} ms")
print(f"1x H100 floor          : {gpu1_floor*1e3:.1f} ms")
print(f"8x H100 TP floor       : {gpu8_floor*1e3:.2f} ms")

The output is 305 chips for the weights alone, 320 KiB of KV cache per token, 1.34 GB (about 5.8 chips of SRAM) for one 4,096-token sequence, a pipeline-only LPU estimate of about 0.88 ms per token, 20.9 ms on one H100 and 2.61 ms on eight. Three lessons fall out of it. First, the single-stream speed advantage is real and comes straight from bandwidth: the estimate is about three times lower than an eight-GPU node, and splitting each stage across chips in parallel lowers it further at the cost of more chip-to-chip traffic. Second, the hardware footprint is large: hundreds of chips across several racks for one model replica. Third, every extra concurrent long-context sequence costs several chips' worth of SRAM, so concurrency and context length, not weights, often decide the real chip count. Change the parameters to your model and context before you believe any vendor comparison.

Using an LPU from software

Most teams meet the LPU through GroqCloud, Groq's hosted API, which is compatible with the OpenAI chat completions format. Groq also publishes a Python SDK. Either works; the OpenAI-compatible route lets you switch an existing client by changing the base URL. Model IDs change as models are added and retired, so read them from the models endpoint rather than hard-coding a name from a blog post.

import os, time
from openai import OpenAI

client = OpenAI(api_key=os.environ["GROQ_API_KEY"],
                base_url="https://api.groq.com/openai/v1")

model = os.environ["GROQ_MODEL"]          # pick one from client.models.list()

t0 = time.perf_counter()
first = None
chunks = 0
stream = client.chat.completions.create(
    model=model,
    messages=[{"role": "user", "content": "Explain KV caching in 3 sentences."}],
    max_tokens=200,
    stream=True,
)
for event in stream:
    if event.choices and event.choices[0].delta.content:
        first = first or time.perf_counter()
        chunks += 1
end = time.perf_counter()
print(f"time to first token {first - t0:.3f}s, "
      f"{chunks / (end - first):.0f} chunks/s after first token")

Measure two numbers separately: time to first token, which includes network, queueing and prompt processing, and inter-token rate after the first token, which reflects the decode path the hardware is optimised for. Streaming chunks are not exactly tokens, so use a non-streaming call and its usage field for exact token counts. Run the same harness against your current provider at the same prompt and output lengths before drawing conclusions; the prefill and decode split article explains why the two phases behave differently.

On the hardware roadmap: in December 2025 a non-exclusive licensing agreement giving Nvidia access to Groq's inference technology was announced, members of Groq's leadership joined Nvidia, and Groq continues to run GroqCloud. Nvidia has since announced LPX racks of 256 Groq 3 LPUs, positioned to handle decode alongside its GPUs. Press coverage of the per-chip specifications is inconsistent, so take figures for newer generations from Nvidia's or Groq's own published specifications rather than from summaries.

Failure modes

  • Assuming fast decode means fast end-to-end. Long prompts are dominated by prefill, and queueing in front of a shared service adds latency the chip never sees. Measure time to first token on your real prompt lengths.
  • Quantisation drift. Fitting weights into SRAM pushes providers toward low-precision formats. Run your own evaluation set against the hosted model rather than trusting the original model's benchmark scores.
  • Context and concurrency limits. SRAM-resident KV cache makes very long contexts expensive. Check the context window and rate limits the endpoint actually offers for each model.
  • Misreading determinism. Deterministic timing is not deterministic output; sampling above temperature zero still varies.

Trade-offs

DimensionLPU (SRAM, static schedule)GPU (HBM, dynamic schedule)
Per-stream decode speedVery high: bandwidth floor far below HBMBounded by HBM bandwidth per GPU group
Memory per chipAbout 230 MB (first generation)80 GB or more
Chips per large modelHundredsOne to eight per replica
Latency varianceFixed per compiled stepVaries with load and contention
FlexibilityCompiler-owned; fixed shapes; no custom kernelsAny model, custom kernels, dynamic shapes
Long context, big batchesCostly: KV competes with weights for SRAMNatural fit: HBM holds large caches
TrainingNot the target workloadPrimary workload

An LPU fits interactive decode on popular models, such as voice agents and agent loops where each step waits on the last. A GPU is safer for custom models, very long contexts, training or offline throughput. Speculative decoding is the GPU-side answer to the same latency problem; benchmark it too.

What to do next

  1. Run the sizing script with your model's layer count, KV heads, precision and target context, and note whether weights or KV cache dominate the footprint.
  2. Build the streaming benchmark against GroqCloud and your current provider, and record time to first token and inter-token rate at your real prompt and output lengths.
  3. Run your own evaluation set on the hosted model to catch quality changes from quantisation before you switch traffic.
  4. List the steps in your product where a user or an agent waits on a sequential LLM call; those are where per-stream decode speed pays off.
  5. Keep a GPU fallback path for models and context lengths the LPU endpoint does not offer, and route by model and length.
  6. Recheck newer-generation specifications from Groq's and Nvidia's own published documents before any capacity plan.
Key takeaway: The LPU trades memory capacity for bandwidth and run-time flexibility for determinism. Weights live in on-chip SRAM, the compiler schedules every cycle, and a large model spans hundreds of chips in a pipeline, which makes single-stream decode very fast and long contexts and custom models expensive. Size your model with the arithmetic, benchmark first-token and inter-token latency separately, check quality on your own evaluations, and route only latency-critical traffic to it.