Every agent framework advertises the same loop: the model reads the conversation, decides to call a tool, the tool runs, its result goes back into the conversation, and this repeats until the model answers. That loop fits in twenty lines. The hard parts are everything around it: a process dying halfway through a payment, a step that needs a person's approval, state that must survive a deploy, and finding out why a run cost forty dollars.

This article takes a framework apart instead of comparing brand names. Product-by-product comparisons live in The Agentic Orchestration Stack Compared. Here you get the six components every framework implements, a small runtime you can run that makes each of them concrete, a worked trace, the failure modes that runtime exposes, and a rubric for deciding whether to adopt a framework, build a thin layer of your own, or combine the two.

Advertisement

What a framework is for

An agent framework is a runtime for a loop whose next step is chosen by a language model. Handing control flow to a model means the next step is not known at compile time, the number of steps is not bounded unless you bound it, and each step can have side effects in the outside world.

That shift creates four obligations. You must describe the actions the model may take, which means tool schemas. You must bound and observe a process whose length you do not choose, which means budgets and traces. You must persist progress, because a run can take minutes or days and outlive the process that started it. And you must decide which actions need a person or a rule to say yes. A framework packages answers to these four obligations plus provider adapters. Judge it on those, not on its hello-world.

The six components

Every framework, whatever it calls things, has the same parts. Naming them gives you a checklist for reading any framework's documentation.

ComponentJobQuestion to ask
Control loopRun step after step until a stop conditionWhat stops it: final answer, step budget, token budget, wall clock?
Model adapterTurn messages and tool specs into a provider request and parse the replyCan I swap providers without rewriting tools and prompts?
Tool layerRegister tools, validate arguments, run them, return results or errorsDo tool errors go back to the model or crash the run?
State and checkpointsHold messages and working data; persist after each stepWhen exactly is a checkpoint taken, and where is it stored?
OrchestrationCompose agents: sub-agents, handoffs, graphs, parallel branchesIs the flow explicit code or emergent from prompts?
Observability and evaluationEmit spans for model and tool calls; score runsCan I replay a failed run and see every prompt and tool result?
The six parts every agent framework implementsCaller / APItask + thread idControl loopstep, budget, stop rulestartModel adaptermessages -> replyprompttool callsTool layerschema, validate, policydispatchTools / MCPAPIs, DBs, codeinvokeApproval gatehuman or ruleside effect?State + checkpointsmessages, step, tokenssave per stepOrchestrationsub-agents, handoffsTraces + evalsspans, cost, scoresemitThe loop is the small part. State, the tool boundary and observability are where frameworks differ and where production bugs live.Read top to bottom: one step = model call, zero or more tool calls, then a checkpoint.
One step is a model call, zero or more tool calls through the tool layer (with an approval gate for side effects), then a checkpoint. Orchestration composes several such loops; traces record every edge.

Frameworks resemble each other most in the loop and adapter, and differ most in state, the tool boundary and observability, which is where incidents come from.

Advertisement

A minimal runtime

The fastest way to understand what a framework does for you is to write the smallest honest version. The code below is plain Python with no dependencies. The model is passed in as a function, so it works with any provider or with a scripted fake for tests. Each piece maps to one row of the table: ToolRegistry is the tool layer, run is the control loop, state plus store.save is state and checkpointing, and approve is the approval gate.

import json

class ToolRegistry:
    def __init__(self):
        self.tools = {}
    def register(self, name, schema, fn, side_effect=False):
        self.tools[name] = dict(schema=schema, fn=fn, side_effect=side_effect)
    def call(self, name, args):
        if name not in self.tools:
            return {"error": f"unknown tool {name}"}
        try:
            return {"ok": self.tools[name]["fn"](**args)}
        except Exception as e:          # errors go back to the model, not up the stack
            return {"error": f"{type(e).__name__}: {e}"}

class Budget(Exception):
    pass

def run(model, registry, state, store, max_steps=8, max_tokens=20000, approve=None):
    while state["step"] < max_steps:
        reply = model(state["messages"], list(registry.tools))
        state["tokens"] += reply["usage"]
        if state["tokens"] > max_tokens:
            raise Budget("token budget exhausted")
        state["messages"].append({"role": "assistant", **reply["msg"]})
        calls = reply["msg"].get("tool_calls", [])
        if not calls:
            store.save(state)
            return reply["msg"]["content"]
        for call in calls:
            spec = registry.tools.get(call["name"], {})
            if spec.get("side_effect") and approve and not approve(call):
                result = {"error": "denied by reviewer"}
            else:
                result = registry.call(call["name"], call["args"])
            state["messages"].append({"role": "tool", "id": call["id"], "content": json.dumps(result)})
        state["step"] += 1
        store.save(state)               # checkpoint after every completed step
    raise Budget("step budget exhausted")

Three design choices matter more than they look. First, tool exceptions are caught and returned to the model as structured errors, because a model that sees KeyError: 'sku' can often correct its own arguments, whereas an exception that escapes the loop kills the run. Second, there are two independent budgets, steps and tokens, because a loop of cheap calls and a short run with an enormous context fail in different ways. Third, a checkpoint is saved after every completed step, so a crashed run can resume from its last step instead of from the start.

Worked example: one refund, traced

To get a deterministic trace, replace the model with a script of three replies: look up the order, request a refund, then answer. The refund tool is marked as a side effect, and the approval rule allows refunds of 50 or less.

class MemoryStore:
    def __init__(self): self.snapshots = []
    def save(self, state): self.snapshots.append(json.dumps(state))

SCRIPT = [   # a scripted model, so the trace is deterministic
    {"msg": {"content": "", "tool_calls": [{"id": "c1", "name": "lookup_order", "args": {"order_id": "A17"}}]}, "usage": 900},
    {"msg": {"content": "", "tool_calls": [{"id": "c2", "name": "refund", "args": {"order_id": "A17", "amount": 40}}]}, "usage": 1100},
    {"msg": {"content": "Refund of 40.00 issued for order A17."}, "usage": 700},
]
def fake_model(messages, tools, _it=iter(SCRIPT)):
    return next(_it)

reg = ToolRegistry()
reg.register("lookup_order", {"order_id": "string"}, lambda order_id: {"id": order_id, "total": 40, "status": "delivered"})
reg.register("refund", {"order_id": "string", "amount": "number"}, lambda order_id, amount: {"refunded": amount}, side_effect=True)
state = {"messages": [{"role": "user", "content": "Refund order A17"}], "step": 0, "tokens": 0}
store = MemoryStore()
print(run(fake_model, reg, state, store, approve=lambda c: c["args"]["amount"] <= 50))
print("steps", state["step"], "tokens", state["tokens"], "checkpoints", len(store.snapshots), "messages", len(state["messages"]))

# Output:
# Refund of 40.00 issued for order A17.
# steps 2 tokens 2700 checkpoints 3 messages 6

Here is how the numbers come about. Step 0: the model asks for lookup_order; the tool has no side effect, so it runs; one assistant message and one tool message are appended; the step counter becomes 1 and the first checkpoint is saved. Step 1: the model asks for refund with amount 40; the tool is a side effect, so approve is consulted and returns true; the refund runs; two more messages; step becomes 2; second checkpoint. Step 2: the model replies with text and no tool calls, so the loop saves a third checkpoint and returns. The total is 2 completed tool steps, 900 + 1,100 + 700 = 2,700 tokens, 3 checkpoints, and 6 messages: the user message, three assistant messages and two tool results.

Change the approval rule to amount <= 30 and the refund tool never runs. The model receives {"error": "denied by reviewer"} as the tool result and must decide what to tell the user. For human approval, approve instead pauses the run, saves state, and resumes when a person decides.

What breaks: the failure modes a framework must handle

The minimal runtime works for the happy path. Each of the following failures is real, and each is a reason frameworks exist. Use this list when evaluating one.

  • Duplicate side effects on resume. The checkpoint is saved after the tools run. If the process dies after the refund succeeds but before store.save, the last checkpoint still says step 1, and resuming calls the model and the refund again. Checkpointing cannot fix this; idempotency can. Derive an idempotency key from data already in the checkpoint and make side-effecting tools honour it, as shown below.
  • Runaway loops. A model that keeps retrying a failing tool will spend the whole budget. Step and token budgets stop the bleeding; detecting repeated identical calls and returning a stronger error stops it sooner.
  • Context growth. Every step appends messages, so cost per step rises with run length. Long runs need summarisation or truncation of old tool results, under your control.
  • Tool output as an attack path. A web page or email returned by a tool can contain instructions. The approval gate protects side effects regardless of why the model asked for them, which is why it belongs in the tool layer and not in the prompt.
  • Schema drift. A tool's signature changes, but saved checkpoints contain calls in the old shape. Version tool schemas and keep old versions resolvable until in-flight runs finish.
  • Parallel calls. Models can request several tools in one reply. Running them concurrently is faster, but results must be appended in a deterministic order or replays will differ.
def refund(order_id, amount, idempotency_key):
    prior = ledger.get(idempotency_key)
    if prior is not None:
        return prior                    # replay after a crash: no second payment
    result = payments.refund(order_id, amount, key=idempotency_key)
    ledger.put(idempotency_key, result)
    return result

# in the loop: derive the key from the thread and the model's call id, both of which are in the checkpoint
args = dict(call["args"], idempotency_key=f"{state['thread']}:{call['id']}")

Long-running runs that must survive restarts and deploys are a workflow-engine problem; Durable Agent Workflows with State, Retries, and Replay goes into replay rules and compensation in depth.

Four control-flow models

Frameworks differ most in how they let you express flow beyond a single loop. There are four models, and most products mix them.

  1. Single loop with tools. One model, many tools, and the model decides everything. Simplest to build and debug; good for tasks under a dozen or so steps with clear tools.
  2. Explicit graph. Nodes are steps, edges are transitions, some edges are chosen by the model. You get testable structure, resumable checkpoints per node and obvious places for approvals. LangGraph for Agent Orchestration shows this model in detail, including the rule that a resumed node re-runs from its start.
  3. Conversation between agents. Several agents with roles exchange messages until a termination rule fires. Flexible, but flow is emergent, so cost and behaviour vary more from run to run.
  4. Hierarchy and handoff. A coordinator delegates sub-tasks to specialist agents, each with a smaller tool set and its own context. This limits context growth and tool confusion, at the cost of extra model calls for routing.

Use the least emergent model that solves the task: put known structure in code, and let the model choose only where the choice depends on content.

The ecosystem moves; design for it

Agent frameworks change quickly. Microsoft, for example, describes Microsoft Agent Framework, which reached 1.0 in 2026, as the successor to both Semantic Kernel and AutoGen, and publishes migration guides from each. Teams that wrote business logic directly against an earlier framework's classes are now migrating. Expect the same churn elsewhere.

To survive churn, keep tools as plain functions with plain schemas behind a thin adapter; the Tools Pattern article describes a framework-neutral tool layer. Keep prompts and policies as data you own, not as framework subclasses. And expose tools through a protocol such as MCP where it fits, so the same tool server works with more than one runtime.

Build, adopt, or both

The minimal runtime above is a real option for narrow agents. Choosing between it and a framework is a trade-off along a few axes.

FactorFavours a thin runtime of your ownFavours adopting a framework
Run lengthSeconds to a few minutes, one processHours or days, must survive restarts
Human approvalRare, or a synchronous ruleFrequent pauses waiting on people
FlowOne loop, a handful of toolsBranches, sub-agents, parallel work
TeamFew people, strong platform skillsMany teams needing a shared convention
ObservabilityYou already have tracing you trustYou need tracing and eval tooling now
Lock-in toleranceLowAcceptable for the features gained

The common middle path is to adopt a framework for state, checkpointing and orchestration, while keeping tools, prompts, approval policies and evaluation datasets in your own code. That way the expensive assets survive a framework change.

Operating agents in production

  • Trace every model call and tool call as a span with inputs, outputs, tokens and latency; Agentic Observability and Tracing covers span design.
  • Set per-run step, token and wall-clock budgets, and alert on the rate of runs that hit them, not only on errors.
  • Store checkpoints in durable storage with a retention policy, and test resume by killing a worker mid-run in staging.
  • Make every side-effecting tool idempotent and route it through a policy check that does not depend on the prompt.
  • Keep an evaluation set of real tasks and run it on every change to prompts, tools, model version or framework version; Agent Evaluation at Scale covers how.
  • Upgrade framework and model versions separately so regressions can be attributed.

What to do next

  1. Run the minimal runtime with the scripted model, then change the approval threshold and the budgets and watch the trace change.
  2. Replace the fake model with a real provider call and keep the same tools; note what you had to add.
  3. Kill the process between a side-effecting tool call and the checkpoint, resume, and confirm your idempotency key prevents a duplicate.
  4. Write down, for one real agent, the answer to each question in the six-component table.
  5. Score your situation against the build-or-adopt table and decide which components you will own.
  6. Move tools, prompts and policies into a framework-neutral layer before your next framework upgrade.
Key takeaway: An agent framework is a runtime for loops whose next step a model chooses. The loop itself is trivial; the value is in the tool boundary, durable state, approval gates, orchestration and observability. Write the small runtime once so you know what each part does, then adopt a framework for the parts that are hard to build well, such as checkpointing and long-running orchestration, and keep tools, prompts, policies and evaluations in code you own. Make side effects idempotent, because checkpointing alone cannot stop a resumed run from repeating them.