Most agent code starts as a loop: call the model, run whatever tool it asks for, append the result, call the model again. That loop works until you need any of the things production asks for: a step that waits two days for a human, a run that survives a deploy, three searches in parallel, a way to see exactly what the agent believed at step seven, or a guarantee that a refund is issued once. LangGraph is a library for writing that loop as an explicit graph whose state is saved after every step, so those requirements become configuration rather than bespoke plumbing.
This article explains how LangGraph runs a graph from first principles, then builds a support-triage agent that classifies a ticket, fans out knowledge-base searches, pauses for human approval before a refund, and resumes safely. It covers the parts that bite in production: reducers and parallel writes, the rule that an interrupted node restarts from its first line, checkpoint growth, concurrent requests on one thread, and when a plain loop is the better choice. API names were checked against the current LangGraph documentation; if your installed version is older, verify signatures before copying code.
The execution model: state, nodes, edges and super-steps
A LangGraph program has three ingredients. The state is a typed dictionary (a TypedDict, dataclass or Pydantic model) that every node can read. A node is an ordinary function that receives the current state and returns a partial update: a dictionary containing only the keys it wants to change. An edge says which node runs next, either unconditionally or by calling a routing function on the state.
Execution proceeds in super-steps, a model borrowed from Google's Pregel graph system. In each super-step, every node scheduled to run executes, possibly in parallel. Their updates are then merged into the state, edges are evaluated to decide which nodes run in the next super-step, and, if a checkpointer is attached, the merged state is saved. A graph ends when no node is scheduled, which normally means the edge led to END.
This has two consequences you should internalise early. First, nodes never mutate shared state directly; they return updates, and the runtime decides how updates combine. Second, the unit of durability is the super-step boundary. If the process dies in the middle of a super-step, the run restarts from the last saved boundary, which means the nodes of the interrupted step run again. Everything about idempotency in LangGraph follows from that.
State schemas and reducers
Each key in the state can carry a reducer, a function that combines the existing value with an incoming update. Without a reducer, an update overwrites the value. With Annotated[list, operator.add], updates are appended. The built-in add_messages reducer, imported from langgraph.graph.message, appends chat messages but replaces a message whose id already exists, which is what you want when a node edits a previous message.
import operator
from typing import Annotated, TypedDict
from langgraph.graph.message import add_messages
class TriageState(TypedDict):
ticket_id: str
messages: Annotated[list, add_messages] # append, replace by id
intent: str # overwrite: last writer wins
kb_hits: Annotated[list, operator.add] # parallel searches append here
order: dict | None
approved: bool | None
draft: strReducers matter most under parallelism. If two nodes run in the same super-step and both return a value for a key with no reducer, the runtime cannot know which one should win and raises an InvalidUpdateError instead of silently picking one. Treat that error as a design signal: either the key needs a reducer, or the two nodes should not both own it. A useful rule is that every key has exactly one writer unless it has a reducer that makes concurrent writes meaningful.
Keep state small and serialisable. Everything in it is written to the checkpointer at every super-step, so a 2 MB retrieved document stored in state is a 2 MB write per step, per thread. Store large blobs elsewhere and keep a reference in state.
Routing: conditional edges, Command and Send
There are three ways to decide what runs next, and choosing well keeps graphs readable.
- Conditional edges.
add_conditional_edges(source, route_fn, path_map)callsroute_fn(state)aftersourcefinishes and maps its return value to a node name. Use these when routing is a pure function of state; the possible destinations are visible in the graph definition, which makes the graph easy to draw and test. - Command. A node can return
Command(update={...}, goto="node")to update state and choose the next node in one step. Use it when the decision is made inside the node's own logic, such as an agent handing off to a specialist, and you would otherwise duplicate that logic in a routing function. - Send. A routing function can return a list of
Send("node", payload)objects. Each one schedules a separate execution of the target node with its own input, all in the same super-step. This is LangGraph's map step; a reducer on the receiving key is the reduce step.
from langgraph.graph import StateGraph, START, END
from langgraph.types import Send
def route_after_classify(state: TriageState) -> str | list[Send]:
if state["intent"] == "refund":
return "lookup_order"
# one search per query, run in parallel in the next super-step
queries = make_queries(state["messages"])
return [Send("search_kb", {"query": q}) for q in queries]
def search_kb(payload: dict) -> dict:
hits = kb_client.search(payload["query"], k=3)
return {"kb_hits": hits} # operator.add merges every branchNote that search_kb receives the Send payload, not the full state. That isolation is useful: each branch sees only what it needs, and the branches cannot interfere with each other except through the reducer.
Worked example: a support-triage agent with approval
The graph in the figure classifies an incoming ticket. Questions fan out to knowledge-base searches and go straight to a drafted reply. Refunds look up the order, draft a reply, then pause for a human to approve before anything is sent.
from langgraph.types import interrupt, Command, RetryPolicy
def classify(state):
intent = llm_classify(state["messages"]) # "refund" or "question"
return {"intent": intent}
def lookup_order(state):
return {"order": orders_api.get(state["ticket_id"])}
def draft_reply(state):
return {"draft": llm_draft(state["messages"], state["kb_hits"], state.get("order"))}
def approve(state):
# Nothing above this line may have side effects: on resume the node restarts here.
decision = interrupt({"ticket": state["ticket_id"], "draft": state["draft"],
"amount": state["order"]["total"]})
return {"approved": decision == "approve"}
def send(state):
if state["intent"] == "refund" and not state["approved"]:
return {}
if state["intent"] == "refund":
payments.refund(state["order"]["id"], idempotency_key=f"refund-{state['ticket_id']}")
mailer.send(state["ticket_id"], state["draft"])
return {}
b = StateGraph(TriageState)
b.add_node("classify", classify)
b.add_node("lookup_order", lookup_order, retry_policy=RetryPolicy(max_attempts=3))
b.add_node("search_kb", search_kb, retry_policy=RetryPolicy(max_attempts=3))
b.add_node("draft_reply", draft_reply)
b.add_node("approve", approve)
b.add_node("send", send)
b.add_edge(START, "classify")
b.add_conditional_edges("classify", route_after_classify, ["lookup_order", "search_kb"])
b.add_edge("lookup_order", "draft_reply")
b.add_edge("search_kb", "draft_reply")
b.add_conditional_edges("draft_reply",
lambda s: "approve" if s["intent"] == "refund" else "send", ["approve", "send"])
b.add_edge("approve", "send")
b.add_edge("send", END)Two details in this code carry most of the safety. The refund call happens in send, a separate node that runs after the interrupt, so it executes exactly once per approved resume, and it still passes an idempotency key to the payment provider in case the process crashes between the provider call and the next checkpoint. And the fan-in from search_kb into draft_reply is a normal edge: the runtime waits until every Send branch of that super-step has finished, then runs the draft node once with all hits merged.
Checkpoints, threads and time travel
Compile the graph with a checkpointer and every super-step is persisted. Runs are grouped by a thread_id passed in the config; reusing a thread id continues that conversation, while a new id starts fresh. InMemorySaver is for tests; for production use the Postgres saver from the langgraph-checkpoint-postgres package, and call setup() once to create its tables.
from langgraph.checkpoint.postgres import PostgresSaver
with PostgresSaver.from_conn_string(DB_URI) as saver:
saver.setup() # run once, e.g. in a migration job
graph = b.compile(checkpointer=saver)
cfg = {"configurable": {"thread_id": "ticket-48213"}, "recursion_limit": 40}
graph.invoke({"ticket_id": "48213", "messages": [user_msg]}, cfg)
snap = graph.get_state(cfg)
print(snap.next) # ("approve",) -> paused, waiting for a human
print(snap.values["draft"])
# later, from the approval UI's request handler:
graph.invoke(Command(resume="approve"), cfg)The snapshot returned by get_state exposes values (the state), next (the nodes that will run next), config (which identifies this exact checkpoint) and tasks. get_state_history(cfg) lists every checkpoint of the thread, newest first. Passing one of those historical configs to invoke replays from that point: nodes before it are not re-executed, nodes after it are. update_state does not roll a thread back; it writes a new checkpoint that branches from the chosen one, leaving the original history intact. That makes it the right tool for an operator correcting a misclassified intent and letting the run continue from there.
Set recursion_limit explicitly. It caps the number of super-steps per run and is your circuit breaker against an agent that loops between two nodes forever. Current documentation gives a default of 1,000 steps from version 1.0.6; older releases used a much smaller default. For a graph like this one, a limit of a few dozen turns a runaway loop into a fast, visible error instead of a large model bill.
The rule that causes most production bugs: a paused node restarts from the top
When you resume after interrupt(), LangGraph does not continue from the line where the interrupt was called. It re-runs the whole node from its first line, and this time interrupt() returns the resume value instead of pausing. The documentation is explicit about this, and its consequences are easy to miss.
Suppose an earlier version of approve created an audit record and then called interrupt(). Every resume would create a second audit record. Worse, if it had called the payment API and then asked for confirmation, the refund would be issued twice. The fix is structural: anything with side effects goes after the interrupt or, better, in a later node.
- Do not wrap
interrupt()in a baretry/except. It pauses by raising a special exception, and a broad except clause swallows the pause. - Do not call
interrupt()conditionally on values that may change between the first run and the resume, and do not reorder multiple interrupts in a node. Resume values are matched to interrupt calls by their position within the node, so a skipped or reordered call hands an answer to the wrong question. - Validate the human's answer in the graph, not in a loop around
interrupt(). Route an invalid answer back to the approve node with a conditional edge so each attempt is its own checkpoint.
The same reasoning applies to crashes. Because a super-step is re-executed after a failure, every node that touches the outside world should either be idempotent or pass an idempotency key derived from stable state, like the ticket id above. A RetryPolicy on a node retries it in-process on exceptions; by default it does not retry common programming errors such as ValueError and TypeError, so a bug fails fast rather than burning three attempts.
Failure modes and how to diagnose them
| Symptom | Likely cause | Fix |
|---|---|---|
| Duplicate emails or charges after approvals | Side effect before interrupt(), re-run on resume | Move it to a later node; add an idempotency key |
| InvalidUpdateError on a parallel step | Two branches wrote a key with no reducer | Add a reducer or give the key one owner |
| Run ends with a recursion limit error | Routing loop, often a validator that never passes | Add a retry counter in state and an exit edge |
| Checkpoint table grows by gigabytes | Large documents in state, long threads never pruned | Store references, prune old checkpoints by thread age |
| Resumed run answers the wrong question | Interrupt order changed between deploys | Never reorder interrupts in a node with live paused threads |
| Two replies sent for one ticket | Two workers invoked the same thread_id concurrently | Serialise per thread with a lock or queue partition |
| Paused threads fail after a deploy | State schema changed under live checkpoints | Add fields as optional; migrate or drain before removing |
Two rows deserve emphasis. Do not assume the library serialises concurrent invocations on one thread: a webhook retry and a user click can both resume the same ticket; serialise per thread yourself. And paused threads are live data: treat graph structure and state schema like a database schema with live rows.
When to use LangGraph, and when not to
LangGraph earns its keep when a run has branches, parallel steps, long pauses, or must survive restarts, and when you need to inspect or fork past runs. It is a poor fit for a single prompt with a tool or two, where a twenty-line loop is easier to read, and for workflows that are long-running business processes with timers, compensation and cross-service guarantees, where a general durable-execution engine is a better foundation. The trade-off is explicit structure in exchange for some ceremony: you write a state schema and edges up front, and you accept the super-step model's rules about idempotency.
For the general theory of durable runs see durable agent workflows with state, retries and replay and agent checkpointing; for designing the states themselves, agent state machine architecture; for scheduling and budgets across many tasks, agent orchestrator architecture; and for how LangGraph compares with other frameworks, the agentic orchestration stack compared.
What to do next
- Write your agent's state as a TypedDict and give every key either exactly one writer or a reducer.
- Draw the graph: name each node, mark which ones call external systems, and make sure those run after any interrupt and pass an idempotency key.
- Build it with InMemorySaver and write tests that interrupt, resume and replay from history.
- Switch to the Postgres saver, run setup() in a migration, and set recursion_limit explicitly on every invocation.
- Add a per-thread lock or queue partition so one thread is never run by two workers at once.
- Add step-level tracing, a checkpoint pruning job, and a rule that interrupt order and required state fields never change while threads are paused.