An agent-to-human handoff happens when an AI agent stops being responsible for a conversation and a person takes over. The customer was talking to software; now they are talking to someone on a support desk, a nurse line or a sales team. Done well, the customer barely notices the seam: the person already knows who they are, what was tried and why the agent stopped. Done badly, the customer waits in silence, repeats everything, and sometimes receives replies from the agent and the human at the same time.

This article treats handoff as an engineering problem with four parts: deciding when to hand off, moving ownership safely, giving the human the right context, and handing the conversation back when that makes sense. It ends with failure modes, the numbers worth tracking and a checklist you can apply to an existing agent.

Advertisement

Handoff is not approval and not collaboration

Three patterns get confused. In an approval gate the agent keeps the conversation and a human approves or rejects one proposed action; the approval gates article covers that. In human-in-the-loop collaboration a person and an agent work on the same task side by side; see human-in-the-loop. In a handoff the owner of the conversation changes. After the transfer the human speaks to the customer and the agent does not, unless it is invited back.

That change of owner is what makes handoff hard. There is now shared state that two parties might write to, a queue in which the customer might wait, and a gap in knowledge between what the agent saw and what the human sees. Every pattern below exists to manage one of those three problems.

The architecture at a glance

Handoff moves ownership of a live conversation, not just a questionCustomer channelchat, voice, emailConversation gatewayowner lock + versionAgent runtimeLLM, tools, memoryHandoff decidertriggers + policyContext packetfacts from logsRouter and queueskills, priority, hoursHuman consolepacket, transcript, toolsEvent logevery owner changeHand-backnotes become constraintsmessagesif owner = agentbuildenqueueassignreturnThe gateway is the single writer: exactly one owner may send to the customer at any moment.Every transition is a compare-and-set on the conversation version and is written to the event log.
Messages pass through a gateway that knows the current owner. The decider triggers a transfer, the packet builder assembles facts from logs, the router assigns a person, and a hand-back path lets the agent resume with the human's notes.

The key component is the conversation gateway. It holds, for each conversation, an owner field (agent, queue or a specific human) and a version number. Every outbound message is checked against the owner before it is sent. Without this single point of control, the agent and the human each believe they are in charge, and the customer sees both.

Around it sit four parts: a decider that evaluates triggers after each agent turn, a packet builder that assembles what the human needs, a router that places the conversation with a suitable person, and a console where that person works. An append-only event log records every change of owner, which you will need for audits and for the metrics later on.

Advertisement

When to hand off: the triggers

Triggers fall into two families. Hard triggers fire immediately and are not weighed against anything: the customer asks for a person, the topic is one your policy reserves for humans (a formal complaint, a bereavement, a legal threat, a safety concern), the needed action has no tool, or the money at stake exceeds what the agent may handle alone. Soft triggers detect that the agent is not making progress: repeated tool failures, the customer restating the same intent, or rising frustration.

SignalExampleMain risk
Explicit requestThe customer types 'talk to a person' or presses a buttonIgnoring it; always honour it
Policy topicA complaint about a fall in a care homeClassifier misses paraphrases
Missing capabilityAddress change, but the tool is disabledAgent improvises instead of stopping
Value thresholdRefund above the autonomous limitLimit set by guesswork
Loop detectionSame intent restated three timesFires too late
Frustration scoreClassifier output above a cutoffFalse positives flood the queue

Encode the triggers as plain code that runs after every agent turn, not as an instruction in the prompt. A prompt can ask the model to suggest a handoff, but the decision should be deterministic and testable. Notice that the frustration score never fires alone: sentiment classifiers are noisy, and a customer who writes in capitals while giving an order number is often perfectly happy.

from dataclasses import dataclass

@dataclass
class Signals:
    asked_for_human: bool        # "agent", "person", "representative" or a button press
    regulated_topic: bool        # e.g. complaint, bereavement, legal threat
    missing_capability: bool     # the needed action has no tool or the tool is disabled
    failed_tool_calls: int       # consecutive failures in this conversation
    repeated_intent: int         # times the same intent was restated
    frustration: float           # 0..1 from a classifier, never used alone
    amount_at_stake: float       # refund or order value, if known

def decide(s: Signals, policy) -> tuple[bool, str]:
    # Hard triggers first: they never wait for a score.
    if s.asked_for_human:
        return True, "customer_request"
    if s.regulated_topic:
        return True, "policy_topic"
    if s.missing_capability:
        return True, "no_capability"
    if s.amount_at_stake > policy.max_autonomous_amount:
        return True, "value_threshold"
    # Soft triggers: the agent is not making progress.
    if s.failed_tool_calls >= policy.max_tool_failures:
        return True, "tool_failures"
    if s.repeated_intent >= policy.max_restatements:
        return True, "loop_detected"
    if s.frustration > policy.frustration_cutoff and s.repeated_intent >= 1:
        return True, "frustration_and_loop"
    return False, ""

The reason string matters as much as the boolean. It travels in the context packet, drives routing, and lets you count which triggers fire, which is how you tune the thresholds later.

Ownership as a state machine

Model each conversation as a small state machine: AGENT_ACTIVE, HANDOFF_PENDING (packet being built), QUEUED, HUMAN_ACTIVE, optionally AGENT_ASSIST (the human owns it, the agent drafts suggestions), and CLOSED. Each transition is a compare-and-set against the stored version, so two processes cannot both move the conversation, and every send checks the owner at the moment of sending.

def transfer(conv_id, expected_owner, new_owner, reason, db):
    row = db.get(conv_id)
    if row.owner != expected_owner:
        return False                      # someone else already moved it
    ok = db.update_if_version(
        conv_id,
        version=row.version,
        owner=new_owner,
        state=STATE_FOR[new_owner],
    )
    if ok:
        db.append_event(conv_id, "owner_changed",
                        frm=expected_owner, to=new_owner, reason=reason)
    return ok

def send_to_customer(conv_id, sender, text, db):
    row = db.get(conv_id)
    if row.owner != sender:               # the agent lost ownership mid-generation
        db.append_event(conv_id, "suppressed_message", sender=sender)
        return
    channel.send(conv_id, text)

The owner check at send time handles the classic race. The agent starts generating a reply, the decider meanwhile triggers a handoff and a human picks up within seconds, and then the agent's reply arrives. Because the owner is no longer the agent, the message is suppressed and logged instead of reaching the customer. Persist this state in the same store as the conversation, and if your agent runs as a long workflow, take a checkpoint at the transition so a restart does not resume the agent's old plan; the checkpointing article explains durable resume.

The context packet: never make the customer repeat

The human needs five things within ten seconds of opening the conversation: who the customer is and whether they are verified, what they want, what has already been done, why the agent stopped, and what the agent would suggest next. Everything else is a link to the full transcript.

{
  "conversation_id": "c_81f2",
  "reason": "value_threshold",
  "customer": {"id": "cus_4410", "verified": true, "verified_by": "otp_sms", "tier": "standard"},
  "intent": "dispute duplicate charge",
  "facts": [
    {"source": "billing.list_charges", "detail": "two charges of 89.00 on 2026-09-28, ids ch_11, ch_12"},
    {"source": "billing.refund_policy", "detail": "duplicate refunds above 50.00 need a human"}
  ],
  "actions_taken": [
    {"tool": "billing.list_charges", "status": "ok"},
    {"tool": "billing.create_refund", "status": "not_called", "why": "above autonomous limit"}
  ],
  "open_questions": ["Was ch_12 a retry by the payment provider?"],
  "suggested_next_step": "refund ch_12 after confirming it is a provider retry",
  "summary": "Customer was charged twice for one order and wants one charge returned.",
  "transcript_ref": "s3://transcripts/c_81f2.jsonl"
}

Build the facts and actions from your tool-call logs, not from the model's account of what it did. A model summary is useful for the one-line summary field, but language models sometimes report that they performed an action they only planned, or round a number. If the human trusts a summary that says 'refund issued' when the log shows no refund call, you get a second refund or none. Keep generated text clearly labelled as a summary and keep factual fields traceable to a source.

Carry the verification state across. If the agent verified the customer with a one-time code, say so and say how; otherwise the human must re-verify, which customers experience as being asked to repeat themselves. Redact what the human does not need: full card numbers and secrets should never be in the packet even if they passed through the conversation.

Routing and the wait

Routing chooses who gets the conversation. Use the trigger reason and intent as skills (billing, clinical, retention), the customer tier and value as priority, and the team calendar as availability. Check availability before promising a human: if nobody with the right skill is on shift, the honest options are a callback, an asynchronous ticket with a stated response time, or an agent that continues with clearly limited scope.

While the customer waits, the agent should not go silent and should not keep solving the problem. A good holding behaviour states the expected wait, offers a callback, and answers only simple questions such as opening hours. The guardrails article describes how to restrict tools by state, which is the mechanism for this reduced mode: in QUEUED the agent loses write tools entirely.

Warm, cold and assisted transfer

StyleHow it worksWhen to use
ColdConversation enters a queue with the packet; first free person takes itHigh volume, short tasks
WarmA specific person accepts first; the customer is told their name before switchingHigh value or upset customers
AssistedHuman owns the conversation; the agent drafts replies and looks things upComplex cases with known tools
ScheduledA callback or ticket replaces a live waitOut of hours, long queues

Assisted mode is the most productive and the most dangerous. The agent's drafts appear only in the console, and the human edits and sends them. Never let a draft auto-send after a timeout; that quietly turns the human back into an approver and reintroduces the double-reply problem.

Handing back to the agent

Many handoffs end with routine work: the human resolves the dispute, and the customer then wants to update a delivery address. Handing back saves human time, but the agent must respect what the human decided. Ask the human for short hand-back notes and store them as constraints attached to the conversation, injected into the agent's context and enforced in code where possible. If the human refused a second refund, the refund tool should reject calls for that order for the rest of the conversation, not merely be told about the refusal in the prompt.

Prevent ping-pong. If a conversation has been handed to humans twice already, a third trigger should go to the same person who handled it before, or stay with a human until it closes.

Worked example: a duplicate charge

  1. The customer writes that they were charged twice. The agent verifies them by SMS code and calls billing.list_charges, which shows two charges of 89.00.
  2. Policy allows the agent to refund duplicates up to 50.00, so the decider fires value_threshold. State moves to HANDOFF_PENDING and the agent's write tools are disabled.
  3. The packet builder pulls the two charge ids and the policy rule from the logs, marks refund as not called, and adds a one-line summary. Routing sends it to billing with standard priority; the estimated wait is four minutes, and the customer is told so.
  4. A billing specialist accepts. The customer sees their first name and sends a message; the specialist confirms ch_12 was a provider retry and refunds it from the console.
  5. The customer asks to change their email address. The specialist hands back with the note 'refund for order 7731 is complete, do not refund again'. The agent updates the email; a later request for another refund on that order is blocked by the constraint.

Failure modes

  • Double replies. Agent and human both answer. Fix with the owner check at send time, not at generation start.
  • Dead air. The handoff fires but nobody is available. Check capacity before promising a person, and alert on queue age.
  • Re-asking. The human asks for details the customer already gave. Audit packets against transcripts weekly.
  • Fabricated history. The summary claims an action that never happened. Source facts from logs and label generated text.
  • Over-triggering. A noisy frustration score sends half the traffic to humans. Require a second signal and review trigger rates.
  • Under-triggering. A policy topic is paraphrased and missed. Test the classifier with adversarial phrasings and keep an always-visible 'talk to a person' option.
  • Data leakage. Secrets copied into the packet or the console. Redact by field type before the packet is stored.

What to measure

Track the handoff rate broken down by trigger reason, time to human (trigger to first human message) at the median and 95th percentile, the re-ask rate (human questions that the packet already answered), resolution after handoff, the bounce rate (conversations handed over more than once), and suppressed messages, which should be rare and tell you how often the race occurs. A falling handoff rate with rising complaints means the agent is holding on to conversations it should release.

What to do next

  1. Add an owner field and version to every conversation, and check the owner inside the send path.
  2. Move handoff triggers out of the prompt into code, with a reason string for each, and unit-test them.
  3. Build the context packet from tool logs; keep the model summary to one labelled line.
  4. Check routing capacity before promising a human, and define the holding behaviour and its reduced tool set.
  5. Store hand-back notes as enforced constraints, and cap repeat handoffs per conversation.
  6. Put the metrics above on one dashboard and review trigger rates every week.
Key takeaway: A handoff transfers ownership, so treat it like any other ownership change: one owner at a time, enforced at the moment of sending, with every transition recorded. Decide with deterministic triggers, give the human facts drawn from logs rather than the model's memory, be honest about the wait, and when the agent takes the conversation back, make the human's decisions binding constraints rather than suggestions.