An LLM agent is a loop. It reads its context, asks the model what to do next, calls a tool, adds the result to its context, and goes round again until it decides it is finished. A kill switch is the mechanism that ends that loop from the outside, on demand, quickly enough to matter, whatever the model happens to be producing at the time.
This sounds easy. In practice many teams find out during an incident that their "stop" button only flips a flag the agent checks between plans, while a ten-minute tool call keeps running. Or the agent's cached cloud token stays valid for another hour. Or the stop request goes through the same message queue that is backed up because of the agent. This article builds a kill switch that holds up under those conditions. It covers the stop levels you need, where each is enforced, the fail-closed lease pattern with working code, automated tripwires, what to do with half-finished work, and how to drill the whole thing. It assumes the agent already runs with least-privilege permissions; a kill switch limits how long damage can go on, and permissions limit how big it can get.
What a kill switch has to guarantee
Write the requirements down before you design anything, because each one rules out a common shortcut.
- Bounded time to stop. Pick a number, for example "no new side effects more than 5 seconds after the stop is issued", and test against it. Without a number, "we stopped it eventually" counts as success.
- Independence from the model. Nothing the model writes can delay, reinterpret or cancel a stop. A stop that is delivered as a message in the agent's context is a suggestion, not a control.
- Fail closed. If the stop machinery itself breaks, for example the control plane is unreachable, agents stop instead of continuing unsupervised.
- Scoped. You can stop one run, every run for one tenant, every run using one tool, or everything, without taking down unrelated work.
- Recoverable state. After a stop you know which actions completed, which were in flight, and which need compensating, so the stop does not turn into a second incident.
Scoping is the point most often skipped. A global off switch is so costly that operators hesitate; a scoped stop gets used early, when it helps most.
Stop levels, from gentle to blunt
No single mechanism meets all five requirements, so build a ladder. Each rung is enforced at a different point and survives the failure of the rung above it.
| Level | Mechanism | Enforced by | Time to effect | Leaves behind |
|---|---|---|---|---|
| 1. Soft cancel | Run marked as cancelling; agent finishes current step and exits cleanly | Agent runtime loop | One step (seconds to minutes) | Clean state, a final checkpoint |
| 2. Hard cancel | Runtime cancels the in-flight model call and tool call | Agent runtime / task supervisor | Sub-second for the loop; tool may not stop | In-flight tool call of unknown outcome |
| 3. Tool deny | Gateway refuses every call for the run's lease | Tool gateway | Next call | Agent may still think and log |
| 4. Credential revoke | Short-lived tokens revoked or not renewed | Credential broker / IdP | Token lifetime or revocation latency | Downstream sessions may persist |
| 5. Isolate | Egress policy cut, container or VM terminated | Network / orchestrator | Seconds | Lost in-memory state |
Levels 1 and 2 depend on the agent runtime behaving correctly. Levels 3 to 5 do not: a stuck loop, a runtime bug or a compromised process is still stopped by the gateway, the identity provider and the network, which the agent cannot reconfigure. The network rung pairs naturally with egress filtering you should already have; the kill switch only needs a way to switch a run's egress profile to "none".
The control plane: out of band and boring
The stop registry is a small, highly available store of stop records. Each record has a scope (run id, tenant, tool, agent version, or global), a reason, who issued it and when. Operators write to it through a CLI or a button, and automated tripwires write through the same API. Enforcement points read from it.
Keep it out of band. It must not share a queue, a database or a rate limit with the agent's normal traffic, because the incidents that need a stop are often the ones that saturate those. A separate small key-value store or a strongly consistent configuration service is enough. Give the write path its own authentication, so that the people who can stop agents are known and logged, and make sure the agent's own identity can only read from it.
Fail-closed leases, with code
The core pattern is a lease. A run may act only while it holds a valid lease from the control plane. The lease is short, for example 15 seconds, and is renewed in the background. Each renewal checks the stop registry. A stop record matching the run means no renewal. An unreachable control plane also means no renewal, and the lease runs out. Either way the run stops within one lease period, with no message to the agent needed.
Enforcement points check the lease locally, so the check is cheap. The runtime checks before every model call and every tool call. The tool gateway checks again on its side, using a signed lease token the runtime attaches to each request, so a buggy runtime cannot skip the check.
import asyncio, time, hmac, hashlib, json, base64
LEASE_SECONDS = 15
RENEW_EVERY = 5
class StopRequested(Exception):
pass
class Lease:
def __init__(self, run_id, scopes, control):
self.run_id, self.scopes, self.control = run_id, scopes, control
self.expires_at = 0.0
self.token = None
async def renew_forever(self):
while True:
try:
# The control plane refuses renewal if any stop record matches a scope.
grant = await asyncio.wait_for(
self.control.renew(self.run_id, self.scopes, LEASE_SECONDS), timeout=2)
self.expires_at, self.token = grant["expires_at"], grant["token"]
except Exception:
pass # do not extend: the lease ages out (fail closed)
await asyncio.sleep(RENEW_EVERY)
def require(self):
if time.time() >= self.expires_at:
raise StopRequested(f"lease expired for {self.run_id}")
return self.token
async def agent_loop(lease, model, tools, state):
while not state.done:
lease.require()
action = await model.next_action(state) # cancellable
token = lease.require() # check again: thinking takes time
result = await tools.call(action, lease_token=token)
state.record(action, result)
def gateway_check(token, secret, now=None):
"""Tool gateway side: verify signature and expiry; never trust the caller's clock."""
body, sig = token.rsplit(".", 1)
good = hmac.new(secret, body.encode(), hashlib.sha256).hexdigest()
if not hmac.compare_digest(sig, good):
return False
claims = json.loads(base64.urlsafe_b64decode(body))
return (now or time.time()) < claims["exp"]Three details matter. Renewal failures are swallowed deliberately, so an unreachable control plane lets the lease expire. The lease is checked again after the model call, which can outlive it. And the gateway verifies the token with its own clock and key, independent of the runtime.
Match the lease length to your time-to-stop target: a 15-second lease allows up to 15 seconds of tool calls after a stop. Shorter leases cost renewal traffic and make control-plane blips stop agents, so set it per agent class.
Cancelling work that is already in flight
A lease stops the next action. The action already running is a separate problem. Model calls are easy: cancel the request and the provider stops generating. Tool calls fall into three groups, and you should label every tool with its group at registration time.
- Cancellable. Read-only queries, searches, local computation. Cancel the task and move on.
- Bounded. Calls with a server-side timeout shorter than your stop target. Let them finish or time out, and record the outcome.
- Unbounded or irreversible. Long jobs, payments, emails, deployments. These need extra care: a hard per-call deadline at the gateway, an idempotency key so a retry after a stop cannot double-apply, and ideally a two-step pattern where the agent only proposes and a separate executor commits after a final lease check.
The two-step pattern is the most useful design change for high-impact tools. The agent writes an intent record ("send this email", "transfer this amount") and returns. An executor process, which holds its own lease and checks the stop registry, picks up intents and performs them. A stop then prevents commits even if the agent has already decided, and the queue of pending intents shows exactly what was about to happen. This also fits naturally with human approval steps for the riskiest actions.
Credentials: the stop that outlives the process
Killing a process does not invalidate what it already holds. A cloud token valid for an hour, or a copied API key, keeps working after the process is gone.
Give agents only short-lived, run-scoped credentials issued by a broker, never long-lived keys. A broker issuing 5- to 15-minute tokens tied to the run id can refuse renewal once a stop exists, and revoke outright where the identity provider supports it. Keep a map from run id to every credential issued, and learn beforehand which downstream systems cache sessions after revocation.
Automated tripwires
People are slow at 3 a.m., so automated detectors should issue stops too. Good tripwires are simple and measured outside the model.
- Budgets. Steps per run, tokens per run, wall-clock time, money spent through tools. Exceeding a hard budget issues a stop for that run. This also covers runaway loops, the agent-level version of LLM denial of service.
- Rates. Tool calls per minute per run or per tenant, especially for write tools. A sudden spike is more often a loop than a productive burst.
- Policy hits. Repeated denials from the tool gateway mean the agent keeps trying something it may not do. Three denials in a row is a reasonable trigger to stop and page someone.
- Repetition. The same tool called with the same arguments several times in a row usually means the agent is stuck.
- Content signals. Detectors on tool output for injected instructions, as in indirect prompt injection, can stop a run before the agent acts on them.
Tune tripwires to stop the run, not the fleet; one that stops everything gets disabled after its first false positive. Escalate from run to tenant to agent version as more runs trip.
Worked example: a support agent issuing refunds
A support agent can read orders and issue refunds of up to $200 through a refunds tool. At 02:10 a prompt-injected product review causes it to start refunding every order it reads. Here is how the layers respond with a 10-second lease and these tripwires: at most 20 refunds per tenant per hour, and a stop for any run with more than 5 refunds.
- 02:10:04 to 02:10:31: the run issues five refunds through the two-step executor, each with an idempotency key.
- 02:10:33: the agent writes a sixth refund intent. The per-run tripwire fires and writes a stop record for the run id.
- 02:10:33: the executor checks the stop registry before committing intent six, finds the stop, and marks the intent cancelled.
- 02:10:38: the lease renewal is refused. The runtime's next
lease.require()raises, and the loop exits. - 02:10:39: the credential broker refuses token renewal for the run; the current token expires at 02:15.
- 02:11: on-call is paged with the run id, the five committed refunds and the cancelled sixth, and the audit log of the review that triggered it.
The tripwire threshold, not a person's reaction time, capped the damage at five refunds, and the recorded intents make reversing them a short script.
After the stop: cleanup and restart
Every stop should produce a run report from the audit log: completed actions, in-flight actions of unknown outcome, and cancelled intents. Resolve the unknowns with the idempotency keys before compensating. Restart deliberately: clear the stop record with a reason, and resume from a checkpoint rather than replaying the context that caused the problem.
Failure modes
- Stop delivered in-band. The stop is a message added to the agent's context. The model may ignore it, summarise it away, or never read it while a tool call is running.
- Fail-open renewal. Lease renewal errors are treated as "keep the old lease", so a control-plane outage turns into unsupervised agents.
- Shared fate. The stop registry lives in the same database or queue the agent is overloading, and the stop request times out.
- Long-lived credentials. The process is dead, but a static API key copied into a sub-process or a scheduled job keeps working.
- Unlabelled tools. A long-running irreversible tool has no deadline or idempotency key, so a stop leaves its outcome unknown and a restart repeats it.
- Global-only switch. The only option is to stop everything, so operators wait too long to use it.
Drills
A kill switch you have not exercised is a guess. Run a canary agent against a test account in production, issue a stop at each level, and measure time to the last side effect against your target. Drill the fail-closed path too, by blocking the control plane and checking that runs stop within one lease period. Repeat after every runtime, gateway or identity change.
What to do next
- Write down your time-to-stop target per agent class, and the scopes you must be able to stop: run, tenant, tool, agent version and global.
- Stand up a small stop registry that is separate from the agent's normal queues and databases, with authenticated writes and audit.
- Add a fail-closed lease checked by the runtime before every model and tool call, and verified independently by the tool gateway.
- Label every tool as cancellable, bounded or irreversible; add deadlines and idempotency keys, and move irreversible tools behind a two-step executor.
- Move agents to short-lived, run-scoped credentials from a broker that stops renewing them when a stop exists.
- Add budget, rate, policy-hit and repetition tripwires that stop the run first and escalate only when several runs trip.
- Run a stop drill at every level this month, measure time to the last side effect, and fix whatever exceeded the target.