Dynamo is the part of torch.compile that turns ordinary PyTorch Python into graphs a compiler can optimise. The overview of the whole stack, Dynamo then AOTAutograd then Inductor, with modes, caching and a first debugging workflow, is in torch.compile, in depth. This page stays inside Dynamo, because almost every surprising behaviour of a compiled model, a slow first step, a recompile storm, a silent fall back to eager or a graph break in an innocent-looking line, is explained by how Dynamo reads your code.
The goal is a mental model that predicts what Dynamo will do with a function before you run it. We follow a call from CPython into Dynamo, watch symbolic interpretation, trace guards to their sources, take a graph break apart, and write a tiny backend that shows exactly what was captured. Log excerpts are quoted from the PyTorch Dynamo deep-dive document; API names and defaults were checked against the PyTorch main branch in October 2026.
The hook: PEP 523 frame evaluation
CPython runs a function by creating a frame and handing it to its evaluation loop. PEP 523, added in Python 3.6, lets an extension replace that evaluation function for an interpreter; it was designed with JIT compilers in mind. When you call a function wrapped by torch.compile, Dynamo installs its own callback, so every frame evaluated inside that call reaches Dynamo before CPython executes it.
The callback does very little on the fast path. Each code object carries a cache of entries, and each entry pairs a set of guards with rewritten bytecode. Dynamo runs the guards of each entry against the frame's actual inputs; the first entry whose guards pass supplies the bytecode, and Dynamo asks CPython's normal evaluation to run it. The bytecode in turn calls compiled graphs. Only when no entry matches does the expensive path start: symbolic tracing, compiling and adding a new entry. Frames from files Dynamo is told to skip, such as much of the standard library, run untouched.
Symbolic interpretation with VariableTrackers
On a miss, an InstructionTranslator walks the frame's bytecode one instruction at a time, with a method for almost every opcode. It does not hold real values on its stack. It holds VariableTracker objects, Dynamo's model of what each Python value is. A tensor becomes a tensor variable carrying an FX proxy and a fake tensor that records dtype, shape, stride and device without real storage. A Python list becomes a list of VariableTrackers. Constants, modules, user functions and many built-ins have their own subclasses. When the error message from a failed trace names a VariableTracker type, it is telling you which of these models Dynamo built for your value.
So Python executes at trace time and disappears: user functions are inlined by a nested translator, loops over containers are unrolled, int and float arithmetic is computed. Only tensor operations become FX graph nodes, with fake tensors propagating metadata so later shape-dependent code can be decided.
import torch
class Block(torch.nn.Module):
def __init__(self, n_layers=3, dim=256):
super().__init__()
self.layers = torch.nn.ModuleList(torch.nn.Linear(dim, dim) for _ in range(n_layers))
self.scale = 0.5
def forward(self, x):
for layer in self.layers: # Python loop: unrolled while tracing
x = torch.relu(layer(x))
if x.shape[-1] > 128: # shape test: decided while tracing, then guarded
x = x * self.scale # Python float attribute: inlined as a constant, guarded
return x
model = torch.compile(Block())
# Schematic of the single FX graph Dynamo records (names simplified):
# l1 = linear(x, w0, b0); r1 = relu(l1)
# l2 = linear(r1, w1, b1); r2 = relu(l2)
# l3 = linear(r2, w2, b2); r3 = relu(l3)
# out = mul(r3, 0.5)
# return (out,)Three things happened to this forward. The loop vanished into three linear and relu pairs. The shape test was evaluated during tracing using the fake tensor, so only one branch was recorded, and a guard now protects that decision. The float attribute became a constant in the graph, protected by a guard on its value. Change self.scale at run time and the guard fails, which is correct and also the first way models end up recompiling.
Sources become guards
Every VariableTracker built from an input remembers its source: where in the frame it came from, such as a local, an attribute of a local, or an item of a list. When tracing depends on a property of a value, Dynamo installs a guard expressed against that source. The deep-dive document's small example shows the format.
# From the PyTorch Dynamo deep-dive document.
@torch.compile
def fn(a, b):
return a * len(b)
fn(torch.arange(10), "Hello") # TORCH_LOGS=guards prints, among others:
# ___check_type_id(L['b'], 94334122025024)
# L['b'] == 'Hello'
fn(torch.arange(10), "Hi") # TORCH_LOGS=recompiles prints:
# Recompiling function fn in script.py:3
# triggered by the following guard failure(s):
# - L['b'] == 'Hello'
# A second deep-dive example, with tensor shapes:
@torch.compile
def fn(a, b):
return a.shape[0] * a * b
fn(torch.randn(4, 3), torch.randn(4, 3))
fn(torch.randn(8, 3), torch.randn(8, 3)) # guards, first call and then the recompile:
# check_tensor(L['a'], torch.float32, device=None, requires_grad=False, size=[4, 3], stride=[3, 1])
# check_tensor(L['a'], torch.float32, device=None, requires_grad=False, size=[None, 3], stride=[3, 1])
# L['b'].size()[0] == L['a'].size()[0]
# 2 <= L['a'].size()[0]Read the guards as assumptions. The trace assumed b was a string equal to "Hello", because len(b) was folded to 5. In the second example, tensor guards check dtype, device, requires_grad, sizes and strides; after the recompile the first dimension is None, meaning symbolic, plus two relational guards: equal first dimensions, and a size of at least 2. Guards let tracing specialise freely because the specialisation is checked on every call; they are also why a Python step counter read in the forward means a new compile every step until the budget runs out.
Shapes as symbols
By default Dynamo first compiles with static shapes. When a guard on a size fails, it recompiles with that dimension as a symbol, a SymInt such as s0, so subsequent sizes reuse the dynamic entry. Three rules shape what you see. Sizes 0 and 1 are always specialised, because emptiness and contiguity depend on them, which is where the 2 <= size guard comes from. Duck shaping assumes that two dynamic sizes equal at trace time are the same symbol, and guards that they stay equal, which lets the compiler fuse across them. And when your code branches on a symbolic size, Dynamo records a guard on the expression; the deep-dive shows one of the form 2*L['a'].size()[0] >= 16, logged with the source line that added it.
You can steer this. torch._dynamo.mark_dynamic(tensor, dim) traces a dimension symbolically from the first compile, avoiding the static-then-dynamic double compile; maybe_mark_dynamic and mark_static are softer and opposite hints. mark_unbacked treats a size as data dependent: it is never assumed to be 0 or 1, assertions about it become runtime checks, and asking for its concrete value raises. Use it when a dimension really can be anything, such as the number of tokens routed to an expert.
Graph breaks are continuation functions
When the translator meets something it cannot model, a call into an unknown C extension, a print, data-dependent Python control flow, it does not give up. It compiles the graph traced so far and rewrites the bytecode into four parts: call the first graph; restore the stack and locals CPython would have had at that point, replaying side effects; run the unsupported instruction normally; and call a resume function that continues from the next instruction.
# Abridged from the deep-dive document (pre-3.11 CPython opcodes; newer Pythons print others).
@torch.compile
def fn(a):
b = a + 2
print("Hi") # unsupported: graph break
return b + a
# MODIFIED BYTECODE fn
# LOAD_GLOBAL __compiled_fn_0 ; graph 1: a + 2
# ... STORE_FAST b ; rebuild the locals CPython would have
# LOAD_GLOBAL print ... CALL_FUNCTION ; run the unsupported call normally
# LOAD_GLOBAL __resume_at_14_1 ; continuation for the rest of fn
# LOAD_FAST a; LOAD_FAST b; CALL_FUNCTION; RETURN_VALUE
#
# MODIFIED BYTECODE resume_in_fn
# LOAD_GLOBAL __compiled_fn_2 ; graph 2: b + a
# LOAD_FAST b; LOAD_FAST a; CALL_FUNCTION; UNPACK_SEQUENCE 1; RETURN_VALUEThe resume function is a new code object with its own frame, so it goes through the PEP 523 hook like any other and is traced, guarded and cached independently. A break in a function called from the compiled function splits the caller's frame as well, which is why one stray print deep in a helper can cut a model into many graphs. Each break also stops fusion across it and, if the unsupported operation needs a value from the GPU, forces a device-to-host synchronisation.
The recompile budget
Every guard failure that leads to a new trace adds an entry to that code object's cache. torch._dynamo.config.recompile_limit caps entries per code object and defaults to 8; cache_size_limit is kept as an alias for older code. accumulated_recompile_limit, 256 by default, caps the total. When a limit is hit, Dynamo warns, names the function and the limit, and stops compiling it: calls that match no existing entry run eagerly, which is how a compiled model becomes slow without failing. TORCH_LOGS=recompiles prints the failing guard for each recompile, which tells you exactly which value keeps changing.
import torch
import torch._dynamo.config as dynamo_config
# Per code object (cache_size_limit is an alias) and across all code objects.
print(dynamo_config.recompile_limit, dynamo_config.accumulated_recompile_limit) # 8 256 by default
x = torch.randn(32, 512, 1024)
torch._dynamo.mark_dynamic(x, 1) # sequence length varies: trace it symbolically from the start
model = torch.compile(model)
for batch in warmup_batches: # one batch per shape you intend to serve
model(batch)
# After warm-up, a recompile in CI is a bug: make it an error.
with torch.compiler.set_stance("fail_on_recompile"):
for batch in eval_batches:
model(batch)
# In serving, run eagerly rather than compile in the middle of traffic.
torch.compiler.set_stance("eager_on_recompile")torch.compiler.set_stance changes behaviour without touching the model. "fail_on_recompile" raises instead of recompiling, turning a silent regression into a test failure. "eager_on_recompile" runs eagerly where a recompile would be needed. "force_eager" ignores torch.compile entirely, and "default" restores normal behaviour.
Writing a backend to see what was captured
A Dynamo backend is a function that receives an FX GraphModule and example inputs and returns a callable. The built-in backends, Inductor, eager and aot_eager, follow the same contract, so a few lines give you a precise view of what Dynamo captured: how many graphs, how large, with what input shapes and which operators.
import torch
from collections import Counter
captured = []
def inspect_backend(gm: torch.fx.GraphModule, example_inputs):
"""A Dynamo backend: receives each captured graph, returns a callable to run it."""
ops = Counter(str(n.target) for n in gm.graph.nodes if n.op == "call_function")
captured.append({
"nodes": len(gm.graph.nodes),
"inputs": [tuple(t.shape) if isinstance(t, torch.Tensor) else t for t in example_inputs],
"top_ops": ops.most_common(3),
})
print(gm.code) # the Python source of the captured graph
return gm.forward # run it eagerly: we only wanted to look
torch._dynamo.reset() # clear caches so earlier compiles do not hide anything
model = torch.compile(Block(), backend=inspect_backend)
model(torch.randn(8, 256))
model(torch.randn(16, 256)) # new batch size: expect a second graph with a symbolic size
for i, g in enumerate(captured):
print(i, g["nodes"], g["inputs"], g["top_ops"])Two graphs where you expected one is a graph break. A second entry with the same structure but different input sizes is a shape recompile. Because this backend runs the graph eagerly, any numerical difference you later see with Inductor belongs to code generation, which kernel fusion and Triton, in depth cover.
Escape hatches, compared
| Mechanism | Dynamo | AOTAutograd and Inductor | Use for |
|---|---|---|---|
torch.compiler.disable | Skips the function; the call site is a graph break | Not involved | Logging, debugging hooks, code you will not make traceable |
torch.compiler.allow_in_graph | Opaque call, not traced inside | Traced inside | Rare; the docs call it a footgun because Dynamo's safety checks are skipped |
torch.compiler.nonstrict_trace | Opaque call, not traced inside | Traced inside | Functions Dynamo cannot handle but AOTAutograd can; pytree inputs, closures treated as constants |
| Custom operator | Opaque | Opaque | Real black boxes such as hand-written kernels |
One more option addresses the commonest break. Setting torch._dynamo.config.capture_scalar_outputs, or TORCHDYNAMO_CAPTURE_SCALAR_OUTPUTS=1, lets Dynamo represent the result of .item() as an unbacked symbol instead of breaking, which helps when the scalar flows into further tensor operations but not when Python branches on it.
Failure modes
- Recompile storm then silent eager: a guarded value, often a Python scalar attribute or global, changes every step until recompile_limit is hit.
- Graph count explodes: a break inside a helper splits every caller; find it with TORCH_LOGS=graph_breaks.
- Shapes 0 or 1 surprise: a batch of size 1 compiles its own specialised entry even with dynamic marks.
- Duck shaping recompiles: two dimensions that matched during tracing later differ, failing the equality guard.
- Stale captures: a closure tensor under nonstrict_trace is a constant, so updates to it are not seen and gradients do not flow to it.
- Long first step: tracing and compiling happen on the first call per entry; warm up before measuring or serving.
Trade-offs
Dynamo's design buys generality: it handles real Python, falls back instead of failing, and stays sound through guards. The costs are guard checks on every call, compile time per entry, and behaviour that depends on values you might not think of as inputs. fullgraph=True trades the fallback for a guarantee and is right once a model is clean. Dynamic shapes trade some kernel specialisation for fewer compiles. Measure on the timeline, as described in GPU profiling, rather than assuming capture succeeded.
What to do next
- Compile your model with the inspection backend above and count graphs and entries for two different batch shapes.
- Run once with TORCH_LOGS=graph_breaks,recompiles,guards and fix the first break and the first recompile reason.
- Move per-step Python scalars that the forward reads into tensors, or out of the forward.
- Mark known-variable dimensions with mark_dynamic before the first call; use mark_unbacked for data-dependent sizes.
- Replace untraceable helpers with torch.compiler.disable for side effects, nonstrict_trace or custom operators for compute.
- Warm up every shape you serve, then wrap evaluation in set_stance("fail_on_recompile") in CI.
- Add fullgraph=True once the model has no breaks, so a future change cannot add one silently.