MTIA, the Meta Training and Inference Accelerator, is Meta's in-house AI chip family. You cannot rent one, so the reason to study it is not procurement. MTIA is the clearest public example of a hyperscaler designing silicon around one dominant workload, recommendation ranking, and then pulling the same design toward generative AI. It is also a case study in making a non-GPU architecture usable from ordinary PyTorch, with Triton as the kernel language.

This article explains the architecture in terms a software engineer can act on: which memory a kernel's bytes come from, why the bandwidth ratios dictate model design, how PyTorch code reaches the chip, and what the 2026 generations change. Every number below comes from Meta's own posts, Meta's research papers, or Meta's Hot Chips 2026 talk as reported by ServeTheHome. Where a detail has not been published, the article says so.

Generations and names

Meta has used two naming schemes, which confuses searches. The first chip, described at ISCA 2023, is now called MTIA 100 (earlier: MTIA v1). The second, announced in April 2024 as the "next-generation MTIA" and described at ISCA 2025 as MTIA 2i, is now MTIA 200. In March 2026 Meta announced four further generations, MTIA 300, 400, 450 and 500, with the stated aim of shipping a new chip every six months or less. Meta says it already runs hundreds of thousands of MTIA chips for inference across organic content and ads.

GenerationPrimary job (per Meta)Status as published
MTIA 100Recommendation inferenceFirst generation, 7nm, LPDDR5
MTIA 200 (2i)Recommendation inferenceDeployed; detailed below
MTIA 300Ranking and recommendation trainingIn production (March 2026)
MTIA 400All workloads, mainly GenAI inferencePresented at Hot Chips 2026
MTIA 450, 500Optimised first for GenAI inferenceAnnounced; few published specs

The workload that shaped the chip

A deep learning recommendation model (DLRM) turns each request into dozens or hundreds of sparse categorical features: pages followed, items viewed, ad IDs. Each feature indexes an embedding table that may hold tens of millions of rows. The model gathers the rows for each feature, pools them (usually a sum), lets the pooled vectors interact, and runs dense MLPs on the result to score candidates.

The gather and pool step is the problem. Summing rows does about one arithmetic operation per element loaded, so at FP16 its arithmetic intensity is about 0.5 operations per byte. A GPU that offers hundreds of teraflops spends that phase waiting on memory, and the tables are too large for its high-bandwidth memory unless they are sharded across many expensive devices. Meta's design response was the opposite of a GPU: modest compute, a large and cheap memory pool, and a big on-chip SRAM to hold the hot part of the working set. The embedding-serving article covers the same problem on GPUs; MTIA is what it looks like when the chip is built for it.

MTIA 200 architecture

Meta's April 2024 post and the ISCA 2025 paper describe MTIA 200 as follows. Compute is an 8×8 grid of processing elements (PEs). Meta's compiler papers describe each PE as a scalar RISC-V core and a vector RISC-V core, plus fixed-function units for operations such as DMA and dot products. Each PE has 384 KB of local memory. The grid shares 256 MB of on-chip SRAM, and the chip has 128 GB of LPDDR5 off-chip.

PropertyMTIA 200 (published)
Process, clock, powerTSMC 5nm, 1.35 GHz, 90 W TDP
INT8 compute354 TOPS dense, 708 TOPS with sparsity
Local memory per PE384 KB
On-chip SRAM256 MB at 2.7 TB/s
Off-chip memory128 GB LPDDR5 at 204.8 GB/s
Host linkPCIe Gen5 x8
RackUp to 72 chips: 3 chassis, 12 boards each, 2 chips per board

Two ratios explain the chip. The first is capacity: 128 GB per 90 W device is far more memory per watt than HBM parts offer, which suits large embedding tables. The second is bandwidth: SRAM is about 13 times faster than LPDDR5. A table row served from SRAM costs a thirteenth of the time of one fetched from LPDDR5, so production performance depends on how much of the hot working set the software keeps on chip. Meta reported 3x per-chip performance, 6x model-serving throughput and 1.5x performance per watt over MTIA 100 on its own models.

MTIA 200 (MTIA 2i): what a kernel sees8 x 8 processing elementseach: scalar + vector RISC-V, fixed-function units,384 KB local memory per PEOn-chip SRAM256 MB, 2.7 TB/sLPDDR5128 GB, 204.8 GB/sHost CPUPCIe Gen5 x8Rack72 chips: 3 x 12 boards x 2Ridge point from LPDDR5: 354 TOPS / 204.8 GB/s = about 1,700 ops per byteRidge point from SRAM: 354 TOPS / 2.7 TB/s = about 130 ops per byteEmbedding-bag sum: about 0.5 ops per byte, so keep hot rows on chip
The memory hierarchy a kernel sees on MTIA 200. The two ridge points show why low-intensity work such as embedding pooling must be served from SRAM.

Roofline arithmetic: where a kernel's bytes come from

The roofline model turns those numbers into decisions. A kernel is compute-bound only if its operations per byte exceed peak compute divided by the bandwidth of the memory it reads. From LPDDR5 that ridge point is about 1,700 INT8 operations per byte; from SRAM it is about 130. An INT8 matrix multiply over a batch of B rows reuses each weight byte B times, giving about 2B operations per byte. So a dense layer whose weights stream from LPDDR5 needs a batch of roughly 850 to stay busy, while one whose weights sit in SRAM needs about 65. Embedding pooling at 0.5 operations per byte never approaches either ridge point. It is bandwidth all the way down.

A five-line calculator makes this concrete for a ranking model. The inputs are the number of tables, lookups per table per sample (pooling factor), embedding dimension, bytes per element and the fraction of lookups served from SRAM:

def embedding_us_per_sample(tables, pooling, dim, bytes_per_elem, sram_hit,
                            sram_bw=2.7e12, dram_bw=204.8e9):
    """Lower bound on embedding-gather time per sample, in microseconds."""
    nbytes = tables * pooling * dim * bytes_per_elem
    seconds = nbytes * (sram_hit / sram_bw + (1 - sram_hit) / dram_bw)
    return nbytes, seconds * 1e6

for hit in (0.0, 0.5, 0.8, 0.95):
    nbytes, us = embedding_us_per_sample(100, 30, 128, 2, hit)
    print(f"hit={hit:.2f}  {nbytes/1e3:.0f} KB/sample  {us:.2f} us  {1e6/us:,.0f} samples/s")

With 100 tables, 30 lookups each and 128-wide FP16 rows, each sample touches 768 KB. With no SRAM hits, LPDDR5 alone allows about 3.75 microseconds per sample, roughly 267,000 samples per second per chip. At an 80 percent hit rate the bound drops to about 0.98 microseconds, about a million samples per second. Caching the hot rows is a 4x lever. Quantising rows to INT8 halves the bytes and doubles both figures. Neither change touches the compute units.

From PyTorch to the chip: Inductor and Triton-MTIA

Engineers at Meta do not write MTIA assembly. Meta describes the stack as PyTorch-first. Models are ordinary PyTorch, torch.compile and TorchInductor lower graphs, and kernels are written in or generated as Triton. The paper "Triton for MTIA" (arXiv 2608.00325, to appear in IEEE Micro) describes a Triton compiler backend for MTIA-2i, changes to Inductor code generation, and minimal language extensions that expose MTIA-specific features. It reports Triton kernels competitive with expert-tuned C++, deployed across about 60 model types and covering 50 percent of layers and 47 percent of non-GEMM execution time. Meta's 2026 announcement names PyTorch, vLLM, Triton and the Open Compute Project as the base.

The public lesson is that a Triton kernel is the portable unit. Here is an ordinary GPU Triton kernel for the embedding-bag sum discussed above. The MTIA extensions themselves are not documented publicly, so this shows the shape of the code, not MTIA-specific calls:

import triton
import triton.language as tl

@triton.jit
def embedding_bag_sum(table_ptr, idx_ptr, offs_ptr, out_ptr, dim,
                      BLOCK_D: tl.constexpr):
    bag = tl.program_id(0)                     # one program per pooled bag
    start = tl.load(offs_ptr + bag)
    end = tl.load(offs_ptr + bag + 1)
    cols = tl.arange(0, BLOCK_D)
    mask = cols < dim
    acc = tl.zeros([BLOCK_D], dtype=tl.float32)
    for i in range(start, end):                # gather: 0.5 ops per byte
        row = tl.load(idx_ptr + i).to(tl.int64)
        acc += tl.load(table_ptr + row * dim + cols, mask=mask, other=0.0).to(tl.float32)
    tl.store(out_ptr + bag * dim + cols, acc, mask=mask)

On a GPU the tuning question is coalescing and occupancy. On a PE grid with explicit local memory and DMA engines, the same kernel's performance depends on how rows are staged from LPDDR5 into SRAM and local memory, which is why Triton needed extensions. If you write Triton for GPUs today (see writing Triton kernels), the habit to carry over is to make memory movement explicit and measurable.

MTIA 300 and 400: HBM, chiplets and a scale-up fabric

Meta's Hot Chips 2026 talk, as reported by ServeTheHome, shows a sharp change of direction. MTIA 300 targets recommendation training. It has 72 PEs and 16 messaging elements, 216 GB of HBM3e, a 3nm compute die and a 5nm I/O die. Moving from LPDDR5 to HBM is the headline: training reads and writes optimiser state and gradients, which needs bandwidth that LPDDR5 cannot supply.

MTIA 400 aims at generative inference. It uses two compute chiplets, each with an 8×6 PE grid plus a redundancy row. RISC-V scalar and vector cores handle general-purpose work. Eight HBM3e stacks provide 9.4 TB/s. An SoC chiplet attaches over PCIe Gen6, and an Ethernet-based scale-up fabric offers 1.2 TB/s and joins up to 72 chips in one domain. The chip has hardware support for MXFP4 plus formats Meta calls MS8 and MS8S, and delivers about 12 PFLOPS at FP4. Against MTIA 200, Meta claims over 15x the FP16 compute and 46x the DRAM bandwidth. Meta says MTIA 450 and 500 are optimised first for GenAI inference, and the talk, per ServeTheHome, pointed to 500 pushing beyond 72-chip domains. Detailed specifications for those two had not been published at the time of writing.

For software, this means the ridge point and the parallelism model both move. LLM decoding is bandwidth-bound like embedding pooling, but over weights and KV cache rather than tables, so HBM bandwidth replaces SRAM hit rate as the main lever. A 72-chip scale-up domain invites tensor and expert parallelism across chips, as NVLink domains do on GPUs. That is presumably why vLLM appears in Meta's stack. The same reasoning applies to TPU v4 and v5 and to HBM-based GPUs.

Failure modes

  • Operator coverage gaps. A model that uses one unsupported operator either fails to compile or falls back to the host, and a PCIe round trip inside a forward pass erases the gains. Meta's own Triton figures (50 percent of layers) show how much engineering coverage takes. Audit graph breaks and fallbacks first.
  • Cache-hit regressions. The roofline above shows that throughput tracks the SRAM hit rate. A new feature with a flat popularity distribution, or a shift in traffic, can cut throughput by several times with no code change. Track hit rate as a production metric.
  • Capacity cliffs. 100 tables of 10 million 128-wide FP16 rows is 256 GB, twice one MTIA 200's LPDDR5. INT8 rows bring it to exactly 128 GB, which leaves no room for anything else. That forces sharding or 4-bit rows, and each choice has its own accuracy or communication cost.
  • Numerics drift across hardware. INT8, FP8 and MXFP4 quantisation and different reduction orders change scores slightly. Ranking models are sensitive to calibration. Shadow-score against the reference platform before moving traffic.
  • Kernel tuning that does not port. Block sizes tuned for a GPU's shared memory are wrong for a PE's 384 KB local memory. Re-tune per target instead of reusing autotune caches.

Trade-offs

DimensionCustom accelerator (MTIA)Merchant GPU
Efficiency on the target workloadHigh; memory system built for itLower for gather-heavy work; strong for dense GEMM
FlexibilityNarrower; new operators need compiler and kernel workBroad; mature libraries for nearly everything
Supply and costControlled by the owner; no vendor marginMarket pricing and allocation
SoftwareOwner must build and maintain the stackVendor and community maintain it
IterationMeta targets a new chip every six months or lessVendor roadmap, typically yearly

A custom chip pays only at a scale where one workload dominates the fleet and a team can own a compiler. For everyone else, the transferable result is a method: measure arithmetic intensity, find which memory each phase reads from, and buy (or rent) bandwidth and capacity in the ratio your model needs. The TPU overview shows Google's version of the same decision.

What to do next

  1. Profile your own ranking or LLM-serving model and split time into gather, dense compute and communication.
  2. Compute arithmetic intensity for each phase and compare it with your hardware's ridge points, as in the calculator above.
  3. Measure the access-frequency distribution of your embedding rows; it decides whether caching or quantisation is the bigger lever.
  4. Move one custom CUDA kernel to Triton and keep a PyTorch reference implementation to check numerics.
  5. Run torch.compile and count graph breaks and fallbacks; that list is your portability debt for any future accelerator.
  6. Re-read Meta's MTIA 400 specs once MTIA 450 and 500 details are published, and redo the ridge-point arithmetic.
Key takeaway: MTIA 200 pairs a modest RISC-V PE grid with 256 MB of SRAM and 128 GB of LPDDR5 because recommendation inference is a memory problem. Performance there is SRAM hit rate. MTIA 300 and 400 move to HBM and a 72-chip scale-up fabric for training and GenAI. The portable lessons are roofline arithmetic, explicit memory movement, and Triton as the kernel layer.