Two agents that have never met have to agree on a surprising number of things before useful work happens: which wire binding to use, which protocol version's semantics apply, how the caller proves who it is, what media types each side can produce and consume, and whether results arrive as a stream, a webhook or by polling. A2A has no handshake message for any of this. The server publishes what it supports in its Agent Card, the client decides, and every request carries the client's decisions so the server can check them.

That makes capability negotiation a client-side algorithm plus a set of server-side refusals. This page builds the algorithm step by step, using field and error names from the A2A 1.0 specification, then shows how to turn each server error into a specific fallback. What the capabilities object means and how servers enforce it is covered in A2A capabilities; this page is about the decision the client makes from it.

Advertisement

What is negotiated, and where each side says it

Negotiation in A2A is declarative and per call. Nothing is agreed once and remembered by the server; the client re-states its choices on every request. That keeps servers stateless with respect to clients, and it means a client can change its mind, for example fall back to an older version, without any renegotiation round.

DimensionServer declares (Agent Card)Client asserts (request)Server's refusal
BindingsupportedInterfaces[].protocolBinding and urlWhich URL and wire format it callsNone: a wrong choice simply fails to connect or parse
VersionsupportedInterfaces[].protocolVersionA2A-Version header or parameterVersionNotSupportedError
CredentialssecuritySchemes, securityRequirements on the card and per skillCredentials in the binding's auth mechanismThe binding's authentication error
Media typesdefaultInputModes, defaultOutputModes, skill inputModes/outputModesPart media types and acceptedOutputModesContentTypeNotSupportedError
Update deliverycapabilities.streaming, capabilities.pushNotificationsWhich method it calls and whether it registers a push configUnsupportedOperationError, PushNotificationNotSupportedError
Extensionscapabilities.extensions[] with requiredA2A-ExtensionsExtensionSupportRequiredError
Capability negotiation: five decisions before the first messageAgent Cardfetched, validatedClient profilewhat this client can do1 Binding + versionsupportedInterfaces2 AuthenticationsecuritySchemes3 Media typesmodes vs parts4 Update deliverystream, push or poll5 Extensionsrequired and optionalCall planurl, headers, configRefuse earlyno viable planall passany failsEach request then carries the plan: A2A-Version, A2A-Extensions, credentials and acceptedOutputModes.The server re-checks every declaration and answers a mismatch with a named error, which maps back to a fallback.
The client compiles a call plan from the card and its own profile, or refuses before sending anything.

Step 1: binding and protocol version

In 1.0 the card lists one or more interfaces in supportedInterfaces. Each names a url, a protocolBinding and the protocolVersion it speaks. The specification describes JSON-RPC, gRPC and HTTP+JSON bindings; use the exact binding strings from the version of the specification you implement rather than guessing case. An agent can expose several interfaces for the same binding at different versions, under the same or different URLs, which is how a server supports old and new clients during a migration.

The client walks its own preference list and picks the first interface whose binding and version it implements. Put the client's preferences first, not the card's order: the client knows what it was tested against. A sensible order is your newest tested version on your preferred binding, then the same version on other bindings, then older versions.

Then send the version on every request in the A2A-Version header, as Major.Minor; patch numbers are not negotiated. The trap is the default: when the header is empty, the specification says the server assumes 0.3. A 1.0 client that forgets the header is silently treated as a 0.3 client, and the failures it sees later look like data errors rather than version errors. Set the header in one place, the transport layer, and test that it is present. How agents evolve capabilities without breaking peers is covered in A2A versioning.

Advertisement

Step 2: credentials

The card's securitySchemes names the authentication schemes the agent understands, such as bearer tokens, OAuth 2.0 flows, API keys or mutual TLS, and securityRequirements says which are required. Skills can carry their own securityRequirements, so a public skill and a privileged skill on the same agent may need different credentials. Resolve requirements per skill, not per agent.

The client picks the first scheme listed for the skill for which it actually holds credentials. If none match, stop: sending an unauthenticated request to learn what will happen leaks intent and burns rate limit. Some agents also publish an authenticated extended card, signalled by capabilities.extendedAgentCard and fetched with GetExtendedAgentCard, which may list skills and modes the public card hides; when that capability is present, negotiate against the extended card after authenticating. Every card field used on this page is described in the Agent Card specification walkthrough.

Step 3: media types

Media types are negotiated in both directions. For input, the parts the client sends must use types the skill accepts: the skill's inputModes when present, otherwise the card's defaultInputModes. For output, the client says what it can consume through acceptedOutputModes in the send configuration, and the agent should tailor its output to that list. When the agent cannot satisfy the request, the specification's answer is ContentTypeNotSupportedError.

Compute both intersections before sending. Treat media types as exact strings unless your own code normalises them; do not assume the server matches wildcards or ignores parameters. Order acceptedOutputModes by preference, structured types first, because a program that can parse application/json should not ask for prose it will then scrape. If the only common output type is one your code cannot process, refuse at planning time.

Step 4: how updates arrive

A2A tasks can run for a long time, so the client must choose how to hear about progress. Streaming keeps a connection open and receives task status and artifact updates as they happen; it requires capabilities.streaming and a client that can hold the connection. Push notifications have the agent call a webhook the client registers; they require capabilities.pushNotifications and a reachable, authenticated endpoint. Polling with GetTask always works, and the returnImmediately flag in the send configuration lets the client get a task handle back without waiting.

The ordering stream, push, poll is a default, not a law. Streaming is the lowest latency but ties a connection to the task's lifetime, which is fragile across proxies and deploys. Push suits tasks measured in hours. Polling is the universal fallback and the right choice for batch jobs. Whatever you choose, keep polling implemented: it is what you fall back to when the other two fail mid-task.

Step 5: extensions

Extensions add behaviour beyond the core protocol and are identified by URI. The card lists them under capabilities.extensions, each with a required flag. The client activates the ones it implements by listing their URIs in the A2A-Extensions header. If the agent marks an extension as required and the client does not declare it, the agent must answer with ExtensionSupportRequiredError. A client should never declare an extension it does not implement just to get past that check; the agent will then send data the client misreads.

The negotiation function

Putting the five steps together gives one pure function from card, skill and client profile to a call plan, or an early refusal. Keeping it pure makes it trivially unit-testable against recorded cards.

from dataclasses import dataclass, field

class NoViablePlan(Exception):
    pass

@dataclass
class ClientProfile:
    bindings: list            # in preference order, e.g. ["GRPC", "JSONRPC"]
    versions: list            # in preference order, e.g. ["1.0", "0.3"]
    credentials: set          # names of schemes we hold credentials for
    accepts: list             # media types we can consume, preference order
    sends: set                # media types we will put in message parts
    can_stream: bool
    push_endpoint: str | None
    extensions: set           # extension URIs we implement

@dataclass
class CallPlan:
    url: str
    binding: str
    version: str
    scheme: str
    accepted_output_modes: list
    delivery: str             # "stream" | "push" | "poll"
    extensions: list = field(default_factory=list)

def negotiate(card, skill_id, me: ClientProfile, schemes_for_skill) -> CallPlan:
    # 1. binding and version: client preference wins, card decides what exists
    iface = None
    for b in me.bindings:
        for v in me.versions:
            iface = next((i for i in card["supportedInterfaces"]
                          if i["protocolBinding"] == b and i["protocolVersion"] == v), None)
            if iface: break
        if iface: break
    if not iface:
        raise NoViablePlan("no shared binding and version")

    # 2. authentication: pick a scheme the skill accepts and we hold credentials for
    scheme = next((s for s in schemes_for_skill if s in me.credentials), None)
    if scheme is None and schemes_for_skill:
        raise NoViablePlan("no usable credentials for this skill")

    # 3. media types: skill modes override card defaults
    skill = next(s for s in card["skills"] if s["id"] == skill_id)
    inputs = set(skill.get("inputModes") or card["defaultInputModes"])
    outputs = skill.get("outputModes") or card["defaultOutputModes"]
    if not me.sends <= inputs:
        raise NoViablePlan(f"agent will not accept {me.sends - inputs}")
    accepted = [m for m in me.accepts if m in outputs]
    if not accepted:
        raise NoViablePlan("no output media type in common")

    # 4. update delivery: stream, else push, else poll
    caps = card.get("capabilities", {})
    if caps.get("streaming") and me.can_stream:
        delivery = "stream"
    elif caps.get("pushNotifications") and me.push_endpoint:
        delivery = "push"
    else:
        delivery = "poll"

    # 5. extensions: every required one must be ours; activate optional ones we know
    exts = caps.get("extensions", [])
    missing = [e["uri"] for e in exts if e.get("required") and e["uri"] not in me.extensions]
    if missing:
        raise NoViablePlan(f"agent requires extensions {missing}")
    active = [e["uri"] for e in exts if e["uri"] in me.extensions]

    return CallPlan(iface["url"], iface["protocolBinding"], iface["protocolVersion"],
                    scheme, accepted, delivery, active)

The transport layer then applies the plan to every request: the URL, the A2A-Version header, credentials for the chosen scheme, the A2A-Extensions header and acceptedOutputModes in the configuration.

POST /a2a/v1 HTTP/1.1
Host: research.example.com
Authorization: Bearer eyJ...
A2A-Version: 1.0
A2A-Extensions: https://example.com/ext/citations/v1
Content-Type: application/json

{"jsonrpc": "2.0", "id": 7, "method": "SendStreamingMessage",
 "params": {"message": {...},
            "configuration": {"acceptedOutputModes": ["application/json"]}}}

Worked example: an orchestrator meets a research agent

A planning orchestrator wants market sizing for a report. Its profile prefers GRPC then JSONRPC, versions 1.0 then 0.3, holds a bearer token, accepts application/json then text/plain, sends text/plain, can stream and implements the citations extension. The research agent's card looks like this:

{
  "name": "Market Research Agent",
  "description": "Produces sourced market sizing reports.",
  "version": "2.4.0",
  "supportedInterfaces": [
    {"url": "https://research.example.com/a2a/v1",
     "protocolBinding": "JSONRPC",  "protocolVersion": "1.0"},
    {"url": "https://research.example.com/a2a/grpc",
     "protocolBinding": "GRPC",     "protocolVersion": "1.0"},
    {"url": "https://research.example.com/a2a/legacy",
     "protocolBinding": "JSONRPC",  "protocolVersion": "0.3"}
  ],
  "capabilities": {
    "streaming": true,
    "pushNotifications": false,
    "extensions": [
      {"uri": "https://example.com/ext/citations/v1", "required": false}
    ]
  },
  "defaultInputModes": ["text/plain", "application/json"],
  "defaultOutputModes": ["text/plain", "application/pdf"],
  "skills": [
    {"id": "market-size", "name": "Market sizing",
     "description": "TAM/SAM/SOM estimate with sources",
     "tags": ["research"],
     "inputModes": ["text/plain"],
     "outputModes": ["application/json", "application/pdf"]}
  ]
}

Step 1 picks the GRPC interface at 1.0, the client's first preference. Step 2 picks bearer. Step 3: the skill overrides the card defaults, so input text/plain is accepted and the output intersection is application/json; text/plain is not offered by this skill, so it drops out. Step 4 picks streaming. Step 5 activates citations, which is optional. The plan is gRPC, version 1.0, bearer, accept application/json, stream, citations active.

Now suppose the orchestrator's gRPC stack is disabled in one environment. The same function picks the JSONRPC 1.0 interface with no other change. Suppose instead the agent's operator removes application/json from the skill: planning fails with no common output type before any network call, and the orchestrator can choose another agent from its registry rather than discovering the problem halfway through a task.

Turning server errors into fallbacks

A card can be stale, cached or wrong, so the server's refusal is the final word. Map each named error to one bounded action, retry at most once with the new plan, and then mark the agent as degraded.

ErrorLikely causeAction
VersionNotSupportedErrorCard changed, or the header was missing or malformedRefetch the card, renegotiate, retry once
ContentTypeNotSupportedErrorPart type or output list not acceptableRecompute media types against the fresh card; convert input or widen the accept list only if your code can handle it
UnsupportedOperationErrorStreaming or another operation not supportedFall back to polling
PushNotificationNotSupportedErrorPush capability absentFall back to streaming or polling
ExtensionSupportRequiredErrorA required extension was not declaredDo not fake it; route to another agent
Authentication failureExpired token or wrong schemeRefresh credentials once, then stop

Cache plans per agent and card version, refresh the card on a schedule, and diff it so changes are deliberate. The lockfile approach in capability discovery pairs well with this: the lockfile records what you tested, and negotiate() refuses to move beyond it without review.

Failure modes

  • Missing version header: a 1.0 client is treated as 0.3 and fails in confusing ways.
  • Card-order selection: the client takes the first interface listed instead of its own tested preference and lands on an untested binding.
  • Agent-level auth: credentials resolved for the agent, not the skill, so privileged skills fail.
  • Accept-anything output: an empty or very wide acceptedOutputModes list returns types the client then cannot parse.
  • Streaming with no fallback: a proxy cuts the stream and the client has no way to recover the task state.
  • Faked extensions: declaring a required extension to pass the check and then misreading its data.
  • Retry loops: renegotiating forever on a persistent error instead of once.

What to do next

  1. Write your client profile down: bindings, versions, schemes, media types, delivery options and extensions, in preference order.
  2. Implement negotiate() as a pure function and unit-test it against recorded cards, including one that should be refused.
  3. Set A2A-Version in the transport layer and add a test that fails if it is absent.
  4. Resolve credentials per skill and fetch the extended card when capabilities.extendedAgentCard is set.
  5. Always send acceptedOutputModes, ordered by preference, structured types first.
  6. Keep polling implemented as the fallback for streaming and push.
  7. Map each named error to one action with a single retry, and alert on agents marked degraded.
Key takeaway: A2A capability negotiation is a client-side decision checked by the server on every call, with no handshake. Read the Agent Card, then choose a binding and version from supportedInterfaces in your own preference order, credentials per skill, media types from the skill's modes, and delivery by stream, push or poll. Activate only extensions you implement. Send the plan on every request, always including the A2A-Version header because an empty header means 0.3. Map each named server error to one bounded fallback.