An agent that works in a demo has shown that the task is possible. An agent in production has to do the task thousands of times a day, inside a cost envelope, on a model its provider may update, using tools that time out, with someone on call for it. The difference is mostly not in the prompt. It is in the operating model around the agent: what you promise, what you version, what you limit, how you release, and what you do at 3 a.m.
This article covers that operating model. Other articles on this site go deep on deployment architecture, tracing, tool reliability, cost and kill switches, and this one links to them where they apply. Here we define service level objectives that fit agents, package everything that changes behaviour into one versioned bundle, enforce budgets inside the loop, release changes through gates, catalogue the incidents agents actually have, and write the runbook.
What makes agents different to operate
A conventional service maps a request to a response through code you wrote, so the same input gives the same output and failures usually throw. An agent is a loop: a model chooses an action, a tool executes it, the result returns to the context, and the loop repeats until the model decides it is done. That produces four properties that drive everything below.
- Variable work per request: one task takes two steps and a similar one takes twenty, so latency and cost are long-tailed distributions, not constants.
- Silent failure: the commonest failure is not an exception but a confident wrong answer or a task abandoned halfway, both returned with HTTP 200.
- Behaviour defined by data: a prompt edit, a tool description or a model update changes behaviour as much as a code change, often without passing the review and CI that guard code.
- Dependencies with their own limits: model providers enforce rate limits and have outages, and tools are other teams' services.
Production readiness is therefore about making those properties visible and bounded: measure outcomes rather than responses, put hard limits on every loop, and treat every behaviour-changing artefact as a release.
The production surface
Five pieces surround the loop. A gateway authenticates callers and applies per-tenant quotas. Runs are queued and executed by workers, because agent runs are jobs that can last minutes. All model calls pass through a model gateway that owns rate limits, retries and fallbacks, so no agent code talks to a provider directly. All tool calls pass through a tool layer that owns timeouts, idempotency and circuit breakers, as described in the tool reliability article. And a bundle registry supplies the versioned configuration each run uses. Traces from every run feed the SLO dashboards and the evaluation service, and an out-of-band kill switch can stop runs at several levels.
Service level objectives for agents
Availability tells you almost nothing about an agent: a run that returns a fluent wrong answer is available. Define SLOs on outcomes, measured from traces and from a graded sample of runs.
| Indicator | How it is measured | Example objective |
|---|---|---|
| Task success rate | Graded sample (human or calibrated judge) plus signals such as a reopened ticket | At least 90% over 28 days |
| Escalation rate | Runs handed to a human or ending in refusal | At most 12% |
| Time to completion, p95 | Trace start to final answer | Under 60 s for interactive tasks |
| Steps per run, p95 | Loop iterations per run | Under 12 |
| Cost per successful task | Model and tool spend divided by successful runs | Under $0.08 |
| Policy violations | Guardrail blocks that reached a tool, plus audit sample | Zero; page on any |
The numbers are illustrations for a support agent, not benchmarks; set yours from a baseline week. Two details matter. Watch cost per successful task rather than cost per run, because a cheaper prompt that fails more often costs more per useful result, as the cost article explains. And keep the grading method stable: if you change the judge prompt, re-baseline, or the SLO moves when nothing in production did.
Error budgets work as for any service. When the success rate over the window falls below objective, freeze behaviour changes other than fixes and spend the time on the failure clusters the traces show.
Version the whole bundle
Behaviour depends on the system prompt, model identifier and sampling parameters, tool definitions, tool service versions, retrieval index snapshot and guardrail policies. Put them in one manifest, hash it, and stamp the hash on every run and span, so you can load exactly last Tuesday's bundle and replay a run.
import hashlib, json
def bundle_id(manifest: dict) -> str:
# Stable id for everything that changes agent behaviour.
canonical = json.dumps(manifest, sort_keys=True, separators=(",", ":"))
return "b-" + hashlib.sha256(canonical.encode()).hexdigest()[:12]
manifest = {
"agent": "support-refunds",
"model": {"id": "provider-model-2026-07-15", "temperature": 0.2, "max_output_tokens": 1024},
"system_prompt_sha": "4be1...e09", # hash of the prompt file in the repo
"tools": {"lookup_order": "v3", "issue_refund": "v5", "send_email": "v2"},
"tool_schemas_sha": "9a0c...d1f",
"retrieval_index": "kb-snapshot-2026-09-28",
"guardrails": "policy-v14",
"budgets": {"max_steps": 12, "max_tokens": 60000, "max_cost_usd": 0.25, "max_seconds": 120},
}
run_bundle = bundle_id(manifest) # stamped on the run record and on every trace spanNote the model entry. Where your provider offers dated or pinned model versions, pin one; an alias that floats to the newest model makes the provider's release your release, without your gates. When a provider retires a version, the upgrade is a bundle change like any other. Fallback models need the same treatment: if the model gateway switches to a different model during an outage, that is a different bundle with its own evaluation results, not a transparent substitute.
Budgets inside the loop
Every loop needs hard limits enforced by the runtime, not requested of the model. A model told to use at most ten steps may still take forty; a runtime that counts will stop it. Four budgets cover most cases: steps, tokens, money and wall-clock time. Add a detector for repeated actions, because the commonest runaway is a model calling the same tool with the same arguments after every failure.
import json, time
def run_agent(task, bundle, model, tools, trace):
b = bundle["budgets"]
start, steps, tokens, cost = time.monotonic(), 0, 0, 0.0
seen = {}
messages = [{"role": "user", "content": task}]
while True:
steps += 1
elapsed = time.monotonic() - start
if (steps > b["max_steps"] or tokens > b["max_tokens"]
or cost > b["max_cost_usd"] or elapsed > b["max_seconds"]):
trace.event("budget_exceeded", steps=steps, tokens=tokens, cost=cost, elapsed=elapsed)
return finish_gracefully(messages, reason="budget") # summarise, hand off
reply = model.complete(messages, tools=tools.schemas(), timeout=b["max_seconds"] - elapsed)
tokens += reply.usage.total_tokens
cost += reply.usage.cost_usd
if reply.final_answer is not None:
return reply.final_answer
messages.append(reply.as_message())
for call in reply.tool_calls:
key = (call.name, json.dumps(call.arguments, sort_keys=True))
seen[key] = seen.get(key, 0) + 1
if seen[key] > 2: # same tool, same arguments, third time
trace.event("repeat_loop", tool=call.name)
return finish_gracefully(messages, reason="loop")
result = tools.execute(call, deadline=start + b["max_seconds"])
messages.append(result.as_message())The model and tool interfaces are illustrative; adapt them to your SDK. The important behaviour is finish_gracefully: when a budget trips, the user gets a summary of what was done and a clean handoff to a person, not a timeout error. Set each limit a little above the 99th percentile of successful runs in your traces, and alert on the budget-trip rate, one of the best early signals of a bad release.
Releasing a change
Treat every bundle change as a deployment, with four gates.
- Offline evaluation: run the new bundle on a versioned task set drawn from production traces, including recent failures, and compare success, steps and cost with the current bundle. Block on regressions beyond a set tolerance.
- Shadow: replay a sample of live traffic against the new bundle with side-effecting tools stubbed, and grade the results. This catches shifts the offline set missed.
- Canary: route a small share of real runs, for example 5%, to the new bundle and compare indicators with the current bundle over the same period. Tasks vary widely, so wait for a few hundred graded runs before reading a success-rate difference of a few points.
- Ramp and bake: raise the share in steps, holding at each until the comparison is stable, and keep the old bundle deployable for instant rollback.
Rollback is a pointer change, which is why bundles must be immutable and every component pinned. Rolling back a prompt but not the tool schema it was written for creates a third, untested bundle.
Dependencies and capacity
Model providers limit requests and tokens per minute and return HTTP 429 when you exceed them. Agents multiply the pressure: one user task becomes ten model calls, and a retry storm during a provider slowdown can turn a partial outage into a full one. The model gateway should hold a per-provider concurrency limit and a token bucket matched to your quota, retry 429 and 5xx responses with exponential backoff and jitter, honour any retry-after header, and shed load by queueing new runs rather than starting runs that will fail midway. A run that fails at step seven has paid for six steps and may have left side effects that need compensation.
Size workers by concurrent runs, not requests per second. By Little's law, runs in flight equal arrival rate times average duration: at 4 runs per second and 25 seconds each, about 100 runs are in flight, each holding a context and calling the model every few seconds. Check that your provider quota covers the token rate those runs generate at peak before launch, not after.
Worked example: a bad release caught by budgets
A refunds agent runs at a 91% success rate and a p95 of 8 steps. A team edits the description of the lookup_order tool to mention a new optional field, and the change ships as a new bundle to a 5% canary. Within two hours the canary shows a budget-trip rate of 6% against 0.4% on the baseline, with repeat-loop events concentrated on lookup_order. Traces show the model now passes the new field with a guessed value, the tool rejects it as invalid, and the model retries with the same guess. The canary success rate is 84% over 380 graded runs.
The comparison triggers an automatic halt and the 5% returns to the old bundle; nobody is paged. The fix has two parts: the tool's validation error now names the field and says to omit it when unknown, and the offline set gains six tasks that reproduce the failure. The rerun canary matches the baseline. Without budgets, those runs would have looped until the provider timeout at several times the cost; without bundle ids on the traces, nobody could have tied the failures to one description edit.
Incident catalogue
| Symptom | Likely cause | First action |
|---|---|---|
| Budget trips and repeat loops rise | Bundle change; tool errors the model cannot act on | Roll back the bundle; inspect repeated calls in traces |
| Success drops with no deploy | Model alias moved, index refreshed, upstream tool changed | Diff component versions over time; pin what floated |
| Latency and 429s climb together | Quota exhausted, retry storm | Cut gateway concurrency, queue new runs, check jitter |
| Cost per task doubles | Verbose tool output inflating context, fallback to a pricier model | Cap tool output size; check fallback rate |
| Duplicate side effects | Retries after timeouts without idempotency keys | Disable the tool, reconcile, add keys |
| Out-of-policy action | Injection through tool output or retrieved content, weak guardrail | Kill switch at the narrowest level, revoke credentials, keep traces |
The on-call runbook
Write the runbook before launch and drill it. A good first fifteen minutes looks like this.
- Stop the harm: if runs are taking out-of-policy actions, use the kill switch at the narrowest level that stops them (one tool, one tenant, one agent), and widen if needed.
- Find the change: compare the bundle ids serving now with an hour ago, and check provider status pages and recent tool deployments.
- Roll back first, diagnose second: if a bundle changed, revert the pointer. If nothing changed on your side, pin the previous model version or disable the affected tool.
- Bound the damage: query traces for affected runs, list their side effects, and start compensation where needed.
- Communicate: tell affected users which tasks may have been handled wrongly. For silent-failure incidents this is the step most often skipped.
- Learn: add the failing tasks to the offline evaluation set, so the same regression cannot pass the first gate again.
Trade-offs
Tighter budgets cut runaway cost and latency but end some legitimate long tasks early; tune them per agent from trace percentiles and revisit them when tasks change. Pinning models buys stability at the price of doing upgrades yourself, on a schedule, before the provider retires the version. A large offline set catches more regressions but slows and costs every change; keep a fast subset for every change and the full set for releases. And autonomy is a dial: an agent that asks approval before every side effect is safe and slow, one that never asks is fast and fragile. Most production agents require approval only for irreversible or high-value actions, enforced by layered guardrails.
What to do next
- Write SLOs on outcomes: success, escalation, p95 time, steps, cost per successful task and policy violations.
- Build the bundle manifest, hash it, and stamp the id on every run and span.
- Pin model versions where the provider allows it, and treat fallback models as separately evaluated bundles.
- Enforce step, token, cost and time budgets plus repeat detection in the runtime, with a graceful finish.
- Send every bundle change through offline evaluation, shadow, canary and ramp, with an automatic halt on regression.
- Route all model calls through a gateway with quotas, jittered backoff and load shedding.
- Write and drill the runbook, including the kill switch, and feed every incident back into the evaluation set.