An agent that receives an underspecified request has three choices: guess, fail, or ask. Guessing books the wrong flight; failing makes the caller start over. The Agent2Agent protocol gives a third option as a first-class task state: the server moves the task to input-required, attaches a question, and waits for a message on the same task that answers it. Done well, this turns a brittle one-shot call into a short conversation that still ends in one task with one result.
Done badly, it produces tasks that wait forever, questions no machine can parse, workers blocked on a user who left, and answers applied to the wrong request. This page is about building the pattern well on both sides: when to ask, how to shape the question, how a server pauses and resumes without holding resources, how a client or a parent agent answers, and how abandoned questions are cleaned up. The full state machine, including a client reducer for every state, is on A2A task lifecycle states; here only the interrupted path is covered. Spec statements below were checked against the A2A 1.0 specification on 2026-10-02.
What the specification fixes
The 1.0 specification names two interrupted states, TASK_STATE_INPUT_REQUIRED and TASK_STATE_AUTH_REQUIRED, distinct from the terminal states completed, failed, canceled and rejected. For input-required it says that the agent requests more input by moving the task to that state, and the client continues by sending a new message with the same taskId and contextId. Its multi-turn example puts the question in the task's status.message, a message with role ROLE_AGENT.
- If a follow-up carries only
taskId, the agent must infer the context; if it carries acontextIdthat differs from the task's, the agent must reject it. - A
taskIdthat does not exist returnsTaskNotFoundError; a message to a task in a terminal state returnsUnsupportedOperationError. - A blocking
SendMessagewaits until the task is terminal or interrupted, so a blocking call returns as soon as the agent asks its question. - The core operation text only requires a stream to close at a terminal state. The HTTP+JSON binding section adds that a stream runs until the task reaches a terminal or interrupted state, then closes, while the auth-required section says agents should keep streams open while waiting for an out-of-band credential.
Because of that last point, a client must handle both behaviours: if the stream stays open, keep reading; if it closes on input-required, send the answer with a new SendStreamingMessage on the same task, or poll with GetTask. The spec does not define the format of the question, how long a server waits, or whether a message that does not answer the question is accepted. Those are the design decisions the rest of this page covers.
When to ask instead of guessing or failing
Asking costs a round trip and, if a human is involved, minutes or hours. Use it when the answer changes the outcome and cannot be derived safely.
| Situation | Move | Why |
|---|---|---|
| A required parameter is missing and has no safe default | input-required | Any guess risks a wrong irreversible action |
| Several valid interpretations with different costs | input-required with options | The caller knows its intent; the agent does not |
| Policy requires explicit confirmation, for example spend above a limit | input-required | Confirmation must be recorded on the task |
| A credential or consent for a downstream system is needed | auth-required | Different handling, often out of band |
| The request is outside the agent's skills | rejected | Asking will not help |
| A parameter is missing but has a documented default | Proceed, and state the default in the result | Avoids needless round trips |
Ask once with everything you need. An agent that asks for the origin, then the destination, then the date costs three round trips where one structured question would do. Collect every missing field before pausing.
Shape the question so a machine can answer it
A text question works for a human in a chat window and fails for a calling agent, which has to guess what is being asked and in what format. Put a human-readable text part and a machine-readable data part in the same status message. In 1.0 a part's type is given by which member is present: text for text and data with a mediaType for structured content.
{
"id": "task-7f3a",
"contextId": "ctx-91",
"status": {
"state": "TASK_STATE_INPUT_REQUIRED",
"message": {
"role": "ROLE_AGENT",
"messageId": "q-7f3a-1",
"parts": [
{"text": "Which cost centre should this expense be charged to?"},
{"mediaType": "application/json",
"data": {
"questionId": "q-7f3a-1",
"expires": "2026-10-03T17:00:00Z",
"answerSchema": {
"type": "object",
"required": ["costCentre"],
"properties": {
"costCentre": {"type": "string", "pattern": "^CC-[0-9]{4}$"}
}
},
"options": ["CC-4410", "CC-4420"]
}}
]
}
}
}The fields inside the data part are a convention you define, not part of the spec. Publish the convention in your agent's documentation or as an A2A extension so callers can rely on it. Include a question ID so answers can be matched, an expiry so callers know how long they have, a JSON Schema for the answer, and options when the set is small. Keep the text part complete on its own so a human-facing client can show it unchanged.
Server side: pause without holding a worker
The worst implementation keeps a coroutine or thread blocked waiting for the answer. It ties up a worker for minutes or days, loses the task when the process restarts, and pins the task to one node. Instead, treat the pause as a checkpoint: save everything needed to continue, publish the state, and release the worker.
def handle_message(msg):
if msg.task_id is None:
return start_new_task(msg)
task = store.get(msg.task_id) # raises TaskNotFoundError
if task.state in TERMINAL:
raise UnsupportedOperationError(task.id)
if msg.context_id and msg.context_id != task.context_id:
raise InvalidParamsError("contextId does not match task")
if store.seen_message(task.id, msg.message_id):
return task # duplicate delivery: no-op
if task.state != "TASK_STATE_INPUT_REQUIRED":
return handle_refinement(task, msg) # server policy, not spec
pending = task.checkpoint.pending_question
answer, errors = parse_and_validate(msg.parts, pending.answer_schema)
if errors:
if pending.attempts + 1 >= MAX_ATTEMPTS:
return store.transition(task, "TASK_STATE_FAILED",
reason="no valid answer")
return store.re_ask(task, pending, errors) # stays INPUT_REQUIRED
# Compare-and-set on the version so two answers cannot both resume.
if not store.cas_state(task, expect_version=task.version,
new_state="TASK_STATE_WORKING"):
return store.get(task.id)
queue.enqueue(resume_executor, task.id, answer)
return store.get(task.id)Three details carry the weight. Deduplicating on messageId makes a retried send harmless, as described in A2A idempotency. The compare-and-set stops two concurrent answers, for example from a human and an automated fallback, from both resuming the task. And the checkpoint holds the plan and partial results, so resuming is a fresh job that reads state, not a long-lived continuation.
Client side: answer from context, policy or a person
A calling client should try the cheapest source of an answer first. Many questions can be answered from the original conversation or a configured policy; only the rest need a person.
def run(client, first_message, resolver, max_rounds=3):
task = client.send_message(first_message) # blocking call
rounds = 0
while task.status.state == "TASK_STATE_INPUT_REQUIRED":
rounds += 1
if rounds > max_rounds:
client.cancel_task(task.id)
raise RuntimeError("agent kept asking; gave up")
question = task.status.message
answer_parts = (resolver.from_context(question)
or resolver.from_policy(question)
or resolver.ask_human(question)) # may take hours
task = client.send_message(Message(
role="ROLE_USER", message_id=new_uuid(),
task_id=task.id, context_id=task.context_id,
parts=answer_parts))
return taskIf a human may take hours, do not keep this loop running. Persist the task ID and question, return control, and resume when the person answers; a push notification webhook, covered in A2A push notifications, tells you when the agent has asked. Always check the state before reading artifacts, because an input-required reply is not a result.
Relaying questions through an orchestrator
In a chain, the agent asking may be two hops away from anyone who can answer. An orchestrator that receives input-required from a sub-agent should first try to answer from its own context, since it often holds the user's original request. If it cannot, it moves its own task to input-required, restates the question for its caller, and records a mapping from its question ID to the sub-task's task ID and question ID.
When its caller answers, the orchestrator validates against its own schema, translates the answer into the sub-agent's schema and sends it on the sub-task. The spec describes exactly this chaining for auth-required; for input-required the same pattern is a natural design rather than a requirement. Two rules keep it safe: rewrite the question in the orchestrator's own terms rather than leaking a sub-agent's internals, and set the orchestrator's expiry shorter than the sub-agent's, so an answer never arrives after the sub-task has already expired. More delegation patterns are covered in A2A orchestration patterns.
Worked example: an expense agent asks for a cost centre
A finance assistant sends an expense agent the text "File my taxi receipt from Tuesday, 42 euros" with a receipt image. The agent extracts amount, date and vendor, but finds the employee belongs to two cost centres. It saves a checkpoint holding the extracted fields, then sets the task to input-required with the question shown earlier, offering CC-4410 and CC-4420 and an expiry of 24 hours.
The assistant has no record of the cost centre, so it asks the employee, who types "engineering". The assistant sends that as a text part on the same task. The agent's validator cannot map free text to the pattern, so it re-asks: the task stays input-required with a new status message, question ID q-7f3a-2 and the error "expected one of CC-4410, CC-4420". The assistant now shows the two options as buttons, the employee picks CC-4410, and the assistant replies with a data part:
{"message": {"role": "ROLE_USER", "messageId": "a-2",
"taskId": "task-7f3a", "contextId": "ctx-91",
"parts": [{"mediaType": "application/json",
"data": {"questionId": "q-7f3a-2", "costCentre": "CC-4410"}}]}}The handler validates it, wins the compare-and-set and enqueues a resume. A worker on another node loads the checkpoint, files the expense and completes the task with an artifact holding the expense ID. The task history shows the original request, both questions and both answers, which is the audit trail finance needs.
Expiry: questions nobody answers
Every pause needs an end. Store an expiry with the question and run a sweeper that acts on tasks past it. Whether expiry means canceled or failed is a server decision; failed with a clear reason in the status message is more informative. Before it fires, a reminder through push notifications can recover many abandoned tasks.
Decide what happens to anything the task holds while paused. A seat hold or inventory reservation should have its own timer no longer than the question's expiry, so it is released even if the sweeper is late. Timer design in general is covered in A2A timeout handling.
input-required versus auth-required
| input-required | auth-required | |
|---|---|---|
| What is missing | Information or a decision | A credential or consent |
| Typical answerer | The caller or its user | An identity provider or user consent flow |
| How it resolves | A message on the same task | A message, or a credential delivered out of band |
| Spec guidance on the stream | May close; the HTTP+JSON binding text says it does | Should be kept open for out-of-band credentials |
| Never do | Ask for secrets in a question | Treat the state change itself as an authorization |
Failure modes
- Free-text-only questions: calling agents guess the format and answer wrongly, causing endless re-asks.
- Blocked workers: a thread waits for an answer, so a restart loses the task and capacity drains during busy hours.
- Double resume: two answers both pass validation and two executors run, creating duplicate side effects.
- Answer on a new task: the client omits
taskId, starting a fresh task while the original waits forever. - Tasks that never expire: thousands of paused tasks hold reservations and storage indefinitely.
- Question loops: an agent re-asks the same question without limit because its validator and its schema disagree.
- Clients reading artifacts early: a caller treats the paused task's partial artifacts as the result.
What to do next
- List the conditions in your agent that should trigger a question, and collect all missing fields into one question.
- Define a data-part convention with question ID, expiry, answer schema and options, and document it for callers.
- Replace any in-process waiting with a checkpoint plus a resume job keyed by task ID.
- Add messageId deduplication and a compare-and-set on the transition out of input-required.
- Cap re-asks per question and fail with a clear reason when the cap is reached.
- Run an expiry sweeper and give every held resource its own shorter timer.
- In clients, handle both stream behaviours and always check the state before reading artifacts.