An agent calls issue_refund(order_id, amount). Thirty seconds later the HTTP client gives up. Did the customer get their money back? The honest answer is that nobody on the agent's side knows. The request may never have left the machine, the payment service may have rejected it, or it may have processed the refund and lost the connection before replying. A naive agent treats the timeout as a failure, the model reads "error" and helpfully calls the tool again, and the customer is refunded twice.

This is the central problem of tool-calling reliability, and it is not solved by retrying harder. It is solved by treating unknown as a real outcome with its own handling, and by building four pieces around it: deadlines that flow from the agent's turn budget down to every socket, idempotency keys that the model cannot accidentally change, an intent log written before each side-effecting call, and a reconciliation step that turns unknown back into known. Compensation then handles the cases where an effect happened and the plan needs it undone.

Error classification, backoff and circuit breakers are covered in agent tool-call recovery architecture. This page covers what the executor between the model and the tool must record so that a call never silently happens twice, or silently not at all.

Advertisement

Three outcomes, not two

Most tool wrappers return success or an exception. For read-only tools that is fine: if a lookup fails, you run it again and nothing in the world has changed. For tools with side effects, a call can end in three states, and they need different handling.

OutcomeWhat you knowTypical signalsCorrect next step
SUCCEEDEDThe effect happened2xx response with a body you parsedRecord the result, continue the plan
FAILED_NO_EFFECTThe effect did not happenValidation error, 4xx before processing, connection refused, DNS failureRetry if transient, or let the model re-plan
UNKNOWNThe effect may or may not have happenedRead timeout, connection reset after the request was sent, 502 or 504 from a proxy, process crash mid-callReconcile before anything else
One tool call: intent first, then the call, then a known outcomeModel proposes calltool + argumentsExecutorkey + deadlineIntent logPENDING, key, undoRemote systemdedupes on key1 write2 callOutcomewithin deadline?SUCCEEDEDeffect happenedFAILED_NO_EFFECTsafe to retry or re-planUNKNOWNtimeout, reset, 5xx after sendresponse okrejectedno answerReconciliation probelook up by key or referencefoundnot foundCompensationregistered before the callif plan abortsResult envelope returned to the model: status, retry_safe, effect summary, messagethe model never guesses whether an effect happened
Every side-effecting call is logged as an intent before it is sent. Its outcome is classified into one of three states, and UNKNOWN is resolved by a reconciliation probe before the model is allowed to act on it.

The dividing line is whether the request could have reached the system that performs the effect. A connection refused or a DNS failure happens before any byte is delivered, so the outcome is FAILED_NO_EFFECT. A read timeout happens after the request was written to the socket, so it is UNKNOWN even though it looks like an error. A 504 from a gateway is UNKNOWN too, because the gateway gave up on the upstream, not the upstream on the work. An UNKNOWN treated as a failure causes duplicate effects; one treated as a success causes silent gaps.

Deadlines: one budget from the turn down to the socket

Agents usually have a budget per user turn, say 60 seconds covering several model and tool calls. If each tool has its own fixed 30-second timeout, the third call can start with 5 seconds of turn left and still wait 30, then complete an effect after the turn was abandoned.

The fix is one absolute deadline, created when the turn starts and passed to every tool call. Each call uses the smaller of its own timeout and the time remaining, and refuses to start when too little is left. Forward the deadline to the remote service too, if it accepts one.

import time

class Deadline:
    def __init__(self, seconds):
        self.at = time.monotonic() + seconds

    def remaining(self):
        return max(0.0, self.at - time.monotonic())

    def timeout_for(self, tool_timeout, reserve=2.0):
        """Per-call timeout: never beyond the turn, and keep a reserve
        so the agent can still report what happened."""
        left = self.remaining() - reserve
        if left < 1.0:
            raise DeadlineExceeded("not enough turn budget to start a side effect")
        return min(tool_timeout, left)

Refusing to start is the important part. A call started with too little time tends to time out after the request was sent, which is UNKNOWN. Refusing before sending keeps the outcome at FAILED_NO_EFFECT, which is cheap to handle.

Advertisement

Idempotency keys the model cannot break

An idempotency key lets the remote system recognise a repeat of a request it already processed and return the original result instead of doing the work again. Payment APIs such as Stripe accept an Idempotency-Key header, and an IETF draft proposes the same header for HTTP APIs generally. How a service stores and replays keys is covered in the idempotency architecture article. The agent side has its own problem: who chooses the key.

If the key is a hash of the arguments, it breaks as soon as the model regenerates them: 49.9 instead of "49.90", an added reason field, a lower-case currency code. Each variation is a new key, so a second refund. If the model supplies the key, it invents a fresh one each time.

The key must come from the executor, derived from the business identity of the action. Each side-effecting tool declares which fields identify it; for a refund, the order and amount, not the free-text reason. The executor canonicalises those fields and combines them with the run id and plan step.

import hashlib, json
from decimal import Decimal

TOOL_IDENTITY = {
    # tool name -> fields that identify the effect, and how to canonicalise them
    "issue_refund": {"order_id": str.upper, "amount": lambda v: str(Decimal(str(v)).quantize(Decimal("0.01")))},
    "send_email":   {"to": str.lower, "template_id": str},
}

def idempotency_key(run_id, step_id, tool, args):
    spec = TOOL_IDENTITY[tool]
    identity = {f: canon(args[f]) for f, canon in spec.items()}
    blob = json.dumps({"run": run_id, "step": step_id, "tool": tool, "id": identity},
                      sort_keys=True, separators=(",", ":"))
    return hashlib.sha256(blob.encode()).hexdigest()[:32]

The step id matters. Two legitimate refunds on one order live in different steps and get different keys; a retry of the same step gets the same key whatever the wording. If the remote service forgets keys after a day and a paused run resumes a week later, the key no longer protects you, and reconciliation has to.

The intent log: write before you call

An idempotency key only helps if the retry uses the same key, and that requires remembering the key across crashes. The executor therefore writes an intent record to durable storage before it sends the request: tool, arguments, key, deadline, the compensation to use if the plan later aborts, and the status PENDING. After the call it updates the record with the outcome. On restart, any record still PENDING is by definition UNKNOWN.

def call_tool(log, run_id, step_id, tool, args, deadline):
    key = idempotency_key(run_id, step_id, tool, args)
    prior = log.get(key)
    if prior and prior.status in ("SUCCEEDED", "FAILED_NO_EFFECT"):
        return prior.envelope()                    # replay the recorded result
    if prior and prior.status in ("PENDING", "UNKNOWN"):
        return reconcile(log, prior, deadline)     # never blind-resend

    log.insert(key=key, run_id=run_id, step_id=step_id, tool=tool, args=args,
               compensation=COMPENSATIONS.get(tool), status="PENDING")
    try:
        resp = TOOLS[tool].invoke(args, idempotency_key=key,
                                  timeout=deadline.timeout_for(TOOLS[tool].timeout))
        log.update(key, status="SUCCEEDED", result=resp)
    except NotSent as e:                            # refused, DNS, pre-flight validation
        log.update(key, status="FAILED_NO_EFFECT", error=str(e))
    except Rejected as e:                           # remote said no, before doing work
        log.update(key, status="FAILED_NO_EFFECT", error=str(e))
    except (ReadTimeout, ConnectionReset, GatewayTimeout) as e:
        log.update(key, status="UNKNOWN", error=str(e))
        return reconcile(log, log.get(key), deadline)
    return log.get(key).envelope()

Map each client library's errors onto NotSent, Rejected and the ambiguous group by asking: could the request have been delivered? When unsure, choose UNKNOWN. A wrong UNKNOWN costs one probe; a wrong NotSent costs a duplicate effect.

A relational table keyed on the idempotency key, indexed on status, works well. If your agent already persists run state, as described in agent checkpointing, write the intent record in the same store so a resumed run sees exactly the calls that were in flight.

Reconciliation: turning unknown into known

Reconciliation asks the remote system what actually happened. Every side-effecting tool should come with a probe: a read-only query that finds the effect if it exists. The best probes look up by the idempotency key or by a client reference you attached to the request, such as a metadata field on the refund. A weaker probe searches by business identity, for example refunds on this order for this amount created after the intent was logged.

def reconcile(log, rec, deadline, attempts=3):
    probe = PROBES.get(rec.tool)
    if probe is None:
        log.update(rec.key, status="UNKNOWN", note="no probe; needs a human")
        return rec.envelope()
    for i in range(attempts):
        if i:                                   # space probes: async writes land late
            time.sleep(min(float(i), deadline.remaining() / 4))
        try:
            found = probe(rec.args, client_ref=rec.key,
                          timeout=deadline.timeout_for(5.0))
        except (ReadTimeout, ConnectionReset):
            continue
        if found is not None:
            log.update(rec.key, status="SUCCEEDED", result=found)
            return log.get(rec.key).envelope()
        if i == attempts - 1:
            # absent after repeated reads: resend with the SAME key
            return resend_same_key(log, rec, deadline)
    return log.get(rec.key).envelope()

Two timing details matter. Some systems apply writes asynchronously, so probe more than once, spaced out, before concluding an effect is absent. And when it is absent, resend with the same key: if the first request is still in flight, the service dedupes the two.

A tool with neither a probe nor idempotency support cannot be made safe automatically. Mark it retry_safe: false and send its UNKNOWNs to a human queue, and ask its owners for a lookup by client reference.

Compensation for a single call

Sometimes an effect happened and the plan needs it undone: the hotel is booked, the flight booking failed, the trip is off. Compensation semantically reverses an effect. Across a workflow this is the saga pattern; for a single call, three rules keep it reliable.

  • Register the compensation with the intent, before the call. If the process crashes after the effect and before the plan aborts, the compensation must still be findable. Store the compensating tool and the fields it needs, such as cancel_booking(booking_ref), where booking_ref is filled in from the result.
  • Compensations are tool calls too. They get their own intent records, keys derived from the original key, deadlines and probes. A compensation can itself time out, and a duplicate cancellation is as possible as a duplicate booking.
  • Some effects cannot be undone. A sent email, a published post or a physical shipment can only be followed by a corrective action, such as a correction email. Put irreversible calls last in a plan, after every step that might fail, and require confirmation for them when the stakes are high.

Never compensate an UNKNOWN call blindly: refunding a charge that never happened is its own incident. Reconcile first, then compensate only confirmed effects.

What the model sees

The model decides the next step from the tool result it reads. If that result is an exception string, it will guess, and its guess under uncertainty is usually to try again. Return a structured envelope instead, with the outcome stated plainly and a flag that tells the model whether trying again is its decision to make.

{
  "status": "UNKNOWN",
  "retry_safe": false,
  "effect": "refund of 49.90 EUR on order A-1182 may or may not have been issued",
  "action_taken": "reconciliation scheduled; operator notified",
  "message": "Do not issue this refund again. Tell the user the refund is being confirmed."
}

Add a system-prompt rule: when retry_safe is false, report the status instead of calling again. The executor enforces this anyway, since a repeated call hits the intent log, but the rule saves wasted turns.

Worked example: a support agent refunds and notifies

A support agent handles "my order arrived broken, please refund it". Its plan has two side-effecting steps: step 3 issues a refund, step 4 emails a confirmation. The turn budget is 45 seconds.

TimeEventIntent logModel sees
t=12sStep 3 starts; key K3 derived from run, step 3, order A-1182, amount 49.90K3 PENDING, compensation: none (refunds are final)-
t=27sRead timeout after the request was sentK3 UNKNOWN-
t=27-31sProbe: list refunds with client reference K3; first read empty, second read finds itK3 SUCCEEDED, refund id re_91SUCCEEDED, refund issued
t=33sModel regenerates step 4 arguments; key K4 from template and address onlyK4 PENDING-
t=35sEmail API returns 202K4 SUCCEEDEDSUCCEEDED, email queued

Without the probe, the agent would have reported a failed refund, and a retry with a fresh key would have refunded twice. Without the intent log, a crash at t=28s would have lost the key. And because the email came last, a truly failed refund would never have been confirmed.

Failure modes and operations

  • Keys from full argument hashes. Model rephrasing breaks them. Derive keys from declared identity fields and the plan step.
  • Timeouts classified as failures. The source of most duplicate effects. Audit each client library's exception mapping and default ambiguous errors to UNKNOWN.
  • Orphaned PENDING records. A crashed executor leaves intents nobody reconciles. Run a sweeper that reconciles any PENDING record older than its deadline.
  • Silent side effects in read tools. A "get" endpoint that also marks something as viewed. Classify tools by what they do.

Monitor, per tool: the share of calls ending UNKNOWN, how they resolved (found, absent, human), compensations run and failed, and stale PENDING records. A rising UNKNOWN rate usually means a timeout sits too close to the service's real latency. Put the idempotency key on every trace span, as described in agentic observability and tracing. The durable-workflow article in this series extends these ideas to whole multi-step runs.

What to do next

  1. List every tool your agent can call and mark each one read-only, idempotent-by-key, or unsafe to retry.
  2. Create one deadline per turn and pass it into every tool call; refuse to start side effects without enough budget.
  3. Declare identity fields for each side-effecting tool and derive keys from run, step and those fields in the executor.
  4. Write an intent record before every side-effecting call and update it with a three-state outcome.
  5. Map each client library's exceptions onto not-sent, rejected and ambiguous, defaulting to ambiguous.
  6. Write a reconciliation probe for each side-effecting tool, and route tools without one to a human queue on UNKNOWN.
  7. Register compensations with the intent, and move irreversible calls to the end of plans.
  8. Return a structured result envelope with a retry_safe flag, and add a sweeper for stale PENDING records.
Key takeaway: A side-effecting tool call can succeed, fail without effect, or end unknown, and reliability comes from handling the third case deliberately. Give each turn one deadline, derive idempotency keys from the action's identity rather than the model's wording, log the intent before calling, reconcile every unknown before acting on it, register compensations up front, and tell the model plainly whether a retry is its call to make.