An A2A agent tells the world what it can do in two different ways. Its skills say what work it performs: reconcile invoices, summarise a contract, book a meeting. Its capabilities say how a client may interact with it while that work happens: whether results can be streamed, whether the agent will call a webhook when a long task changes state, whether an authenticated caller can fetch a fuller card, and which protocol extensions it understands. Skills are about the job. Capabilities are about the plumbing, and getting the plumbing wrong produces errors that look like network faults.
This article treats the capabilities object as a contract with two sides: the client reads it to plan how it will talk to the agent, and the server must enforce exactly what it advertised. Facts are checked against the A2A 1.0 specification at a2a-protocol.org, read on 2026-10-01. The full field list of the card is covered in the Agent Card specification, field by field; this page is about behaviour.
What the capabilities object is
In A2A 1.0 the Agent Card carries a required capabilities object of type AgentCapabilities. It has three optional booleans and one optional list:
| Field | What it promises | Operations it gates |
|---|---|---|
streaming | The agent can deliver task updates over a live stream | SendStreamingMessage, SubscribeToTask |
pushNotifications | The agent can call a client-supplied webhook when a task changes | CreateTaskPushNotificationConfig and the Get, List and Delete config operations |
extendedAgentCard | An authenticated caller can fetch a fuller card | GetExtendedAgentCard |
extensions | The agent understands these protocol extensions, each identified by a URI | Any operation, when an extension is marked required |
Two consequences follow from the shape. First, the object is required even when every flag is false, so a client never has to guess whether the field was forgotten. Second, a flag that is absent means the same as false. A client must not read an absent streaming as unknown and try anyway; the specification says the streaming operations MUST fail in that case.
Capabilities are also distinct from the protocol version. The version on each entry of supportedInterfaces says which message shapes and method names the endpoint speaks; capabilities say which optional features are switched on within that version. A 1.0 endpoint with streaming off is a perfectly valid 1.0 endpoint. Version negotiation is covered in A2A protocol versioning; this page assumes both sides already agree on 1.0. Note too that the 0.3 flag stateTransitionHistory is no longer part of AgentCapabilities in 1.0, so code that branches on it is reading a field that will never appear.
Server-side enforcement: the errors the specification names
Advertising a capability is a promise; not advertising one is also a promise, that the operation will be refused cleanly. The A2A 1.0 specification names the refusals:
- If
capabilities.streamingis false or not present,SendStreamingMessageandSubscribeToTaskMUST returnUnsupportedOperationError. - If push notifications are not supported, the push notification configuration operations return
PushNotificationNotSupportedError. - If the extended card is not offered,
GetExtendedAgentCardreturnsUnsupportedOperationError; if the capability is declared but no extended card has been configured,ExtendedAgentCardNotConfiguredError. - If an extension is marked
required: trueand the client did not declare support for it, the agent MUST returnExtensionSupportRequiredError.
The important engineering point is that these checks belong in one place, before any handler runs, and that they read the same configuration that produced the card. The most common way to violate the contract is to hand-edit the card in one repository while the feature flags live in another. A gate generated from the card source cannot drift from it.
from dataclasses import dataclass, field
@dataclass(frozen=True)
class Capabilities:
streaming: bool = False
push_notifications: bool = False
extended_agent_card: bool = False
extended_card_configured: bool = False
extensions: dict = field(default_factory=dict) # uri -> {"required": bool}
def to_card(self):
return {
"streaming": self.streaming,
"pushNotifications": self.push_notifications,
"extendedAgentCard": self.extended_agent_card,
"extensions": [{"uri": u, "required": e["required"]}
for u, e in self.extensions.items()],
}
GATED = {
"SendStreamingMessage": ("streaming", "UnsupportedOperationError"),
"SubscribeToTask": ("streaming", "UnsupportedOperationError"),
"CreateTaskPushNotificationConfig": ("push_notifications", "PushNotificationNotSupportedError"),
"GetTaskPushNotificationConfig": ("push_notifications", "PushNotificationNotSupportedError"),
"ListTaskPushNotificationConfigs": ("push_notifications", "PushNotificationNotSupportedError"),
"DeleteTaskPushNotificationConfig": ("push_notifications", "PushNotificationNotSupportedError"),
"GetExtendedAgentCard": ("extended_agent_card", "UnsupportedOperationError"),
}
def gate(caps, method, requested_extensions):
"""Return an error name, or None if the call may proceed."""
if method in GATED:
flag, error = GATED[method]
if not getattr(caps, flag):
return error
if method == "GetExtendedAgentCard" and not caps.extended_card_configured:
return "ExtendedAgentCardNotConfiguredError"
for uri, ext in caps.extensions.items():
if ext["required"] and uri not in requested_extensions:
return "ExtensionSupportRequiredError"
return NoneMap the returned error name to the binding's error representation in one function, and log the method and the missing capability as structured fields. A refusal is not an outage, so count it separately from server errors; a sudden rise in UnsupportedOperationError almost always means a client is reading a stale card.
Extensions: declared on the card, activated per request
Extensions let an agent add behaviour that the core protocol does not define, such as an extra metadata convention or a domain-specific message part, without forking the protocol. Each entry in capabilities.extensions is an AgentExtension with a uri that identifies it, a human-readable description, a required flag and optional params that configure it for this agent. Version the URI itself, for example by ending it in /v1, so that an incompatible revision is a different extension rather than a silent change.
Declaring an extension on the card does not switch it on for every call. The extensions topic page describes activation per request: the client sends an A2A-Extensions header whose value is a comma-separated list of extension URIs it wants to use, the agent ignores requested extensions it does not support, and the response SHOULD carry an A2A-Extensions header listing those that were actually activated. A client should therefore read the response header rather than assume that asking was enough.
POST /a2a HTTP/1.1
Content-Type: application/json
A2A-Version: 1.0
A2A-Extensions: https://example.com/ext/cost-report/v1, https://example.com/ext/trace-tags/v2
HTTP/1.1 200 OK
A2A-Extensions: https://example.com/ext/cost-report/v1Here the client asked for two extensions and the agent activated one; the client must not expect trace tags in the response. The required flag inverts the default. A required extension is one the agent cannot operate without, typically because it changes the meaning of the messages, and a client that does not activate it gets ExtensionSupportRequiredError on every call. Use required sparingly. Every required extension is a hard dependency imposed on every client, and a client written before the extension existed can no longer talk to you at all.
The client side: plan from the card, fall back cleanly
A well-behaved client turns the capabilities object into a delivery plan before it sends the first message. For a task that may run for minutes, the preference order is usually a live stream, then a webhook, then polling. Each step down costs latency or load but never correctness, because GetTask works against every agent.
import random, time
def plan_delivery(card, have_webhook_endpoint):
caps = card.get("capabilities", {})
if caps.get("streaming") is True:
return "stream"
if caps.get("pushNotifications") is True and have_webhook_endpoint:
return "push"
return "poll"
def wait_for_task(client, card, task_id, deadline_s=900):
mode = plan_delivery(card, client.webhook_url is not None)
if mode == "stream":
try:
for event in client.subscribe_to_task(task_id):
client.handle(event)
if client.is_terminal(event):
return client.get_task(task_id)
except client.UnsupportedOperation:
client.refresh_card() # the card was stale; fall through to polling
except client.StreamDropped:
pass # resubscribe is also fine; polling is the floor
if mode == "push":
client.create_push_config(task_id, client.webhook_url)
delay, start = 1.0, time.monotonic()
while time.monotonic() - start < deadline_s:
task = client.get_task(task_id)
if client.is_terminal(task):
return task
time.sleep(delay + random.uniform(0, delay / 2))
delay = min(delay * 2, 30.0)
raise TimeoutError(task_id)Three details matter. The client compares with is True rather than truthiness, so an absent flag is treated as false exactly as the specification intends. Push mode still polls, because webhooks can be lost and the poll loop is the reconciliation path; with push working it simply finds the task finished on its first or second check. And a refusal from a gated operation triggers a card refresh, because it is the strongest evidence the client has that its cached card no longer matches the deployment. How streams and webhooks behave once chosen is covered in A2A streaming architecture and A2A push notification architecture.
Worked example: one agent, three deployments
Consider a document-review agent run by one team in three places. In the public region it sits behind a gateway that buffers responses, so streaming is unreliable there. In the internal cluster streaming works and a webhook dispatcher exists. A partner gets a dedicated deployment with an extended card that lists two extra skills.
| Deployment | streaming | pushNotifications | extendedAgentCard | Client plan |
|---|---|---|---|---|
| Public region | false | true | false | push, with polling as reconciliation |
| Internal cluster | true | true | false | stream, polling on failure |
| Partner | true | false | true | stream; fetch the extended card after authenticating |
The team first published one card for all three, copied from the internal cluster. Public-region clients chose streaming, the buffering gateway held events until the task finished, and users saw a frozen progress bar. Nothing returned an error. The fix had two parts: generate each deployment's card from that deployment's configuration, and make the gate refuse SendStreamingMessage in the public region so a stale client gets a clear UnsupportedOperationError instead of a degraded stream.
The partner case shows the extended card's role. Its public card advertises extendedAgentCard: true; after authenticating, the partner calls GetExtendedAgentCard and receives a card that the specification says may contain additional details or skills not present in the public card. The specification does not say capabilities may differ between the two cards, so the safe design is to keep capabilities identical and put only skills and descriptive detail behind authentication.
Capability drift and how to catch it
Drift is the gap between what a card claims and what the running deployment does. It appears in predictable places:
- Infrastructure changes under the agent. A new proxy, CDN rule or API gateway starts buffering, and streaming degrades without the card changing.
- Feature flags toggled independently. The webhook dispatcher is turned off during an incident and nobody updates the card.
- Partial rollouts. During a deploy, half the replicas support an extension and half do not, while the card advertises it.
- Cached cards. Clients hold a card for hours; the agent changed its capabilities ten minutes ago.
Defend in layers. Generate the card and the gate from one source, so the card can never claim what the gate refuses. Bump the card's version and its ETag whenever capabilities change, so caching clients revalidate. Roll out a new capability by enabling it on every replica before advertising it, and remove one by withdrawing it from the card first, waiting longer than the card's cache lifetime, and only then disabling it. That ordering means a client holding a stale card meets a working feature, never a missing one.
Testing the contract
Capabilities are easy to test because the expected behaviour is binary. Run a conformance suite against every deployment, in CI and as a synthetic check in production, that fetches the live card and calls every gated operation:
def test_capabilities_match_behaviour(client):
card = client.fetch_card()
caps = card["capabilities"]
probes = {
"streaming": ("SubscribeToTask", "UnsupportedOperationError"), # dummy id: no side effect
"pushNotifications": ("ListTaskPushNotificationConfigs", "PushNotificationNotSupportedError"),
"extendedAgentCard": ("GetExtendedAgentCard", "UnsupportedOperationError"),
}
for flag, (method, refusal) in probes.items():
result = client.call(method, client.minimal_params(method))
if caps.get(flag) is True:
assert result.error_name != refusal, f"{flag} advertised but {method} refused"
else:
assert result.error_name == refusal, f"{flag} not advertised but {method} answered"
for ext in caps.get("extensions", []):
if ext.get("required"):
r = client.call("GetTask", {"id": "probe"}, extensions=[])
assert r.error_name == "ExtensionSupportRequiredError"For streaming, go one step further and measure the time to the first event against the time to completion on a task that emits several updates. If they are nearly equal, something between client and agent is buffering, and the flag is true on paper only.
Trade-offs
| Decision | Benefit | Cost |
|---|---|---|
| Advertise streaming | Lowest latency to first result, natural progress UX | Long-lived connections through every proxy; harder load balancing |
| Advertise push only | No held connections; works across firewalls in one direction | Client must expose and secure a webhook; delivery can be lost |
| Advertise neither | Simplest to operate and scale | Clients poll; latency equals the poll interval |
| Required extension | Guaranteed semantics on every call | Locks out every client that has not adopted it |
| Optional extension | Gradual adoption | Agent must handle both shapes indefinitely |
Failure modes
- Truthiness bugs. A client that treats a missing flag as maybe and tries anyway gets refusals it does not expect.
- Silent degradation. Streaming advertised through a buffering proxy returns no error, just useless latency.
- Assuming activation. A client that sends
A2A-Extensionsand never reads the response header processes output as if every extension were active. - Required-extension lockout. Marking an extension required in a minor release breaks every existing client at once.
- Withdraw-after-disable. Turning a capability off before removing it from the card guarantees refusals for every client holding the old card.
What to do next
- Generate each deployment's card and its capability gate from one configuration object, and delete any hand-maintained card.
- Implement the gate before any handler, returning the error names the 1.0 specification requires, and log refusals as their own metric.
- In every client, plan delivery from the card with strict boolean checks, keep polling as the floor, and refresh the card on any refusal.
- Read the
A2A-Extensionsresponse header and branch on what was activated, not what was requested. - Add the conformance probe to CI and as a production synthetic check, including a time-to-first-event measurement for streaming.
- Write down the order for adding and removing a capability, enable before advertise and withdraw before disable, and follow it in every rollout.