A single request to a remote agent is easy: send a message, get a task back, read its artifacts. Real work is rarely one request. The remote agent needs a missing detail, the user changes their mind, a result prompts a follow-up, or two specialist agents must be coordinated over many turns. That is a conversation, and conversations are where agent systems go wrong: they lose context, ask the same question twice, loop between agents, or never end.

The A2A protocol gives you the building blocks, contexts, tasks and messages, but not the architecture that uses them across many turns and many agents. This article describes that architecture for an orchestrating agent that talks to several remote agents on a user's behalf. It assumes you know the Message object and its identifier rules from A2A messages and builds the conversation layer on top: mapping, continuation, history, budgets, termination and operations. Names follow A2A 1.0.

Advertisement

What a conversation is in A2A

Three identifiers carry conversation structure. A contextId groups related tasks and messages into one logical conversation with one agent. A taskId identifies one unit of work inside that context, with its own lifecycle; the task state machine covers it in detail. A messageId identifies one turn. The specification's rules, section 3.4, are short: an agent MAY generate a new contextId when a message arrives without one, task ids are always minted by the server, an agent MUST infer the context from the task when only taskId is given, and it MUST reject a message whose contextId and taskId do not match.

Two consequences shape everything that follows. First, the remote agent owns its identifiers. Your orchestrator cannot choose the contextId for a new conversation with a remote agent it has never talked to; it learns it from the first response. Second, context is per agent. A conversation that spans a flights agent and a hotels agent is, at the protocol level, two separate contexts that know nothing about each other. Joining them is your job.

The orchestrator's conversation store

The orchestrator keeps one record per user conversation, and inside it the mapping to each remote agent's context and the tasks it has created there. This store, not any remote agent, is the source of truth for what the user sees, what was promised and what is still open.

from dataclasses import dataclass, field

@dataclass
class RemoteThread:
    agent_url: str
    context_id: str | None = None          # learned from the agent's first reply
    open_task_id: str | None = None        # task awaiting input, if any
    task_ids: list[str] = field(default_factory=list)
    last_state: str | None = None

@dataclass
class Conversation:
    conversation_id: str                   # ours, e.g. the user session id
    threads: dict[str, RemoteThread] = field(default_factory=dict)   # by agent name
    turns_used: int = 0
    deadline: float = 0.0                  # absolute, monotonic seconds
    summary: str = ""                      # compact state shared with agents

Persist this record transactionally after every remote reply, before acting on it. If the orchestrator crashes mid-conversation, it must be able to reload the record, see which tasks are open, and resume by querying those tasks rather than by re-sending messages and creating duplicates.

One user conversation, several remote A2A contextsUsersession s-91Orchestrator agentturn manager + budgetsConversation stores-91 maps to remote contextsflights: ctx-F1, hotels: ctx-H7Flights agentcontextId ctx-F1task t1 COMPLETED, t3 WORKINGHotels agentcontextId ctx-H7task t2 INPUT_REQUIREDSendMessageSendMessageRemote task historyGetTask / ListTasks, historyLengthrehydrateTermination rulesmax turns, deadline, loop detector, terminal statesTrace: conversation id s-91 plus every messageId, taskId and contextIdEach remote agent mints its own contextId; the orchestrator owns the mapping, the budgets and the user-facing transcript.
The orchestrator maps its own session id onto a context per remote agent, tracks open tasks, rehydrates history from the remote side when needed, and enforces termination rules across all of them.
Advertisement

Continue a task or start a new one

Every outgoing turn requires one decision: does this message continue an existing task, or start a new task in the same context? The task's state decides.

Remote task stateNext message shouldIdentifiers to send
TASK_STATE_INPUT_REQUIREDAnswer the agent's question in the same tasktaskId (and the matching contextId)
TASK_STATE_AUTH_REQUIREDComplete authorisation, then continue the same tasktaskId
TASK_STATE_WORKING or SUBMITTEDUsually wait or subscribe; send more input only if the agent supports ittaskId
TASK_STATE_COMPLETED, _FAILED, _CANCELED, _REJECTEDStart a new task for any follow-upcontextId, referenceTaskIds with the earlier task
No task yetStart the conversationNo ids; store the contextId the agent returns

The terminal row matters most. The specification says tasks in a terminal state cannot accept further messages, and the server answers such a message with an unsupported-operation error. So a refinement such as make it a window seat after a completed booking is a new task in the same context, and the specification says clients SHOULD reference the related task through referenceTaskIds. That gives the remote agent an explicit pointer to the earlier result instead of relying on it to guess from context.

def next_message(thread, text, ref_task=None):
    msg = {"messageId": new_uuid(), "role": "ROLE_USER",
           "parts": [{"text": text}]}
    if thread.open_task_id and thread.last_state in (
            "TASK_STATE_INPUT_REQUIRED", "TASK_STATE_AUTH_REQUIRED"):
        msg["taskId"] = thread.open_task_id
        msg["contextId"] = thread.context_id
    elif thread.context_id:
        msg["contextId"] = thread.context_id        # new task, same conversation
        if ref_task:
            msg["referenceTaskIds"] = [ref_task]
    return msg                                       # no ids: agent starts a context

History: what to send and what to fetch

A tempting mistake is to resend the entire transcript on every turn, as you would with a stateless chat model. In A2A the remote agent already keeps task history inside its context. Resending it duplicates tokens, risks contradicting the agent's own record and leaks more data than the agent needs. Send only the new turn, plus a short summary when the agent needs facts from other agents.

When the orchestrator itself needs the remote side's view, after a restart or to show the user what happened, it fetches it. GetTask accepts a historyLength parameter that limits how many of the most recent messages are returned, and ListTasks can filter by contextId, with paging and its own historyLength, to enumerate every task in one conversation. SendMessage configuration also carries a historyLength for the response. Ask for little: the last few messages are usually enough to rebuild state, and large history payloads cost latency on every hop.

Sharing context between agents

Because contexts are per agent, information from one agent reaches another only through the orchestrator. Decide explicitly what crosses. A good default is a compact, structured summary of decided facts, such as the destination, dates, traveller count and the chosen flight's arrival time, rather than raw transcripts. Large results travel as artifact references the next agent can fetch, not as pasted content.

This is also a privacy boundary. The hotels agent does not need the passport number the flights agent asked for. Minimising what crosses each boundary reduces both token cost and the blast radius if one agent is compromised or misbehaves. Record in the conversation store which facts were shared with which agent, so you can answer later what any agent was told.

Turn budgets and termination

Conversations between agents have no natural end. Two agents can politely ask each other for clarification forever, or an orchestrator can keep retrying a task that will never succeed. Termination must be designed, with several independent stops.

  • Turn budget. A maximum number of outgoing messages per conversation and per remote task. When reached, escalate to the user rather than continuing.
  • Deadline. An absolute time budget for the whole conversation, propagated into per-call timeouts; see timeout handling.
  • Loop detection. If the remote agent asks a question that is semantically the same as one already answered, do not answer it again automatically; surface it.
  • Terminal states. COMPLETED, FAILED, CANCELED and REJECTED end a task. Close the corresponding open entry and decide whether the conversation as a whole is done.
  • Explicit cancellation. When the user abandons the goal, cancel open remote tasks with CancelTask instead of leaving them working, as described in task cancellation.

Make the conversation's own end state explicit too: done, abandoned, escalated or expired. A conversation without an end state is one nobody can clean up.

Worked example: a trip across two agents

A user asks the orchestrator to book a flight to Lisbon on 12 March and a hotel near the arrival airport for three nights. The turn budget is twelve messages and the deadline five minutes.

  1. Turn 1, to flights with no ids. The agent returns task t1 in context ctx-F1 in TASK_STATE_INPUT_REQUIRED, asking for the departure city. The orchestrator stores ctx-F1 and t1 as open.
  2. The orchestrator already knows the user's home airport from the session, so turn 2 answers in t1 with taskId and contextId set. The agent completes t1 with an itinerary artifact; t1 is closed.
  3. Turn 3, to hotels with no ids, carrying a summary: Lisbon, arriving 12 March 18:40, three nights, near the airport. The agent returns t2 in ctx-H7 in INPUT_REQUIRED, asking for a budget.
  4. The orchestrator has no budget on file, so it asks the user, who says under 150 euros a night. Turn 4 answers in t2; the agent completes it with two options.
  5. The user then asks for a later flight. t1 is terminal, so turn 5 goes to flights with contextId ctx-F1, no taskId, and referenceTaskIds [t1]. The agent creates t3, aware of the earlier itinerary.
  6. The new arrival time changes nothing for the hotel, so no message is sent to hotels. Five turns used; both threads closed; the conversation is marked done.

Note what did not happen: no transcript resent, no question asked twice, no message sent to a terminal task, and no hotel re-query that the change did not require.

Concurrency, retries and expiry

Two turns for the same task can race, for instance a user double-click or an orchestrator retry after a timeout. Serialise outgoing messages per remote task in the orchestrator, and reuse the same messageId when retrying a send whose outcome is unknown, since agents MAY use it to detect duplicates; idempotency in A2A covers the details. Before retrying, query the task: if the earlier send created or advanced it, there is nothing to retry.

Contexts do not live forever. The specification lets agents implement context expiration or cleanup and says they SHOULD document the policy. Read the policy for each agent you depend on, store when each context was last used, and treat a not-found error for a context or task as a signal to start a fresh context with a summary, not as a fatal failure. Long-running user sessions should expect this routinely.

Observability

Debugging a multi-agent conversation without correlation is guesswork. Give each user conversation a trace id and attach every messageId, taskId and contextId to spans as attributes, so a single query shows the whole tree of turns across agents. Log each turn's direction, state transition and token cost. Useful metrics are turns per completed conversation, the share of conversations ending in each end state, input-required rounds per task, and the rate of rejected messages to terminal tasks, which should be zero and indicates a state-tracking bug when it is not.

Failure modes

  • Context confusion. Sending a taskId from one agent with a contextId from another, which the spec requires agents to reject. Keep ids per thread, never in shared variables.
  • Orphaned tasks. The user leaves and remote tasks keep working or waiting. Cancel on abandonment and sweep open tasks past the deadline.
  • Lost turns after a crash. The orchestrator replies to the user before persisting the remote reply. Persist first, then act.
  • Context drift. Summaries passed between agents grow stale after the user changes a fact. Rebuild the summary from the store on every cross-agent turn.
  • Silent expiry. A long-idle context has been cleaned up remotely. Detect not-found and restart with a summary.

Trade-offs

Relying on remote history keeps messages small but makes you dependent on each agent's retention policy; keeping a full local transcript gives independence at the cost of storage and a second copy to secure. Strict turn budgets prevent runaway conversations but occasionally stop a legitimate long negotiation, which is why the budget should escalate rather than fail. Sharing richer summaries improves agent accuracy but widens privacy exposure. One context per remote agent for the whole session is simple; a new context per user goal isolates unrelated tasks better but loses continuity. Choose per agent, and write the choice down.

What to do next

  1. Define a conversation record that maps your session id to a per-agent context, open task and state, and persist it after every reply.
  2. Implement the continue-or-new decision from task state, and use referenceTaskIds for every follow-up to a terminal task.
  3. Stop resending transcripts; send new turns plus a structured summary, and fetch remote history with a small historyLength when needed.
  4. Add a turn budget, a deadline, loop detection and cancellation on abandonment, each escalating to the user.
  5. Serialise sends per task, reuse messageId on retries, and query before retrying.
  6. Document each remote agent's context expiry policy and handle not-found by restarting with a summary.
  7. Trace every turn with conversation, context, task and message ids, and alert on messages sent to terminal tasks.
Key takeaway: In A2A a conversation is a context owned by each remote agent, made of tasks and messages, so an orchestrator talking to several agents must own the mapping between its user session and those contexts. Continue a task only while it is waiting for input, start a new task with referenceTaskIds after it ends, send new turns and summaries rather than transcripts, fetch history sparingly with historyLength, and enforce turn budgets, deadlines and cancellation so every conversation reaches a recorded end state.