An A2A agent advertises what it can do in its Agent Card, and client agents decide whether and how to call it from that advertisement. Most guidance stops at the card: which fields exist and how to fill them in. The harder engineering lives behind it. Someone has to decide which skills a given caller may see, which internal handler a request is sent to when the protocol does not say which skill the caller wanted, what happens when a caller tries something it was never shown, and how a capability is retired without breaking the agents that depend on it.
This article covers that server-side capability layer. It separates the three things the word capability is used for, proposes one internal registry from which the public card, the extended card, the request router and the execution policy are all derived, and works through routing, enforcement, extension activation and lifecycle with code. The card's fields are covered field by field in the A2A Agent Card specification; this page assumes them. Statements about the protocol refer to the A2A 1.0 specification as published at a2a-protocol.org.
Three things called capability
Conversations about agent capabilities go wrong because three different layers share one word. Keeping them apart is most of the design.
| Layer | Where it appears | Who consumes it | Changes when |
|---|---|---|---|
| Protocol capabilities | capabilities in the card: streaming, push notifications, extended card, extensions | Client SDKs, before choosing an operation | You deploy new protocol features |
| Skills | skills in the card: id, name, description, tags, examples, media types | Routing models, registries, human integrators | You add, change or retire a business function |
| Internal capabilities | Nowhere public: tools, model endpoints, data sources and the permissions to use them | Your own runtime and policy engine | Every release |
Protocol capabilities are switches: each flag gates an operation, and a client must check it before trying. Skills are descriptions written for routing: they tell a caller what to ask for, not how the agent does it. Internal capabilities are implementation detail and should never leak into the card. One skill such as reconcile an invoice may use five internal tools; one tool may serve several skills.
One registry, four projections
The failure pattern to avoid is the hand-written card: a JSON file edited by whoever shipped the last feature, drifting away from what the code actually supports. A skill that is advertised but not routable produces confused callers; a skill that is routable but not advertised is an undocumented API that will be found and depended on.
Instead, keep one internal registry of capability manifests and generate everything from it. The public card is the projection of manifests marked public. The extended card is the projection a specific authenticated caller is entitled to. The router's index is the set of handlers that caller may reach. The execution policy is the set of scopes and limits each handler enforces. Because all four read the same records, they cannot disagree.
The capability manifest
A manifest records the skill as it should be advertised plus everything the server needs to route and enforce it. A minimal one looks like this:
{
"id": "reconcile-invoice",
"state": "ga",
"visibility": "extended",
"skill": {
"name": "Reconcile an invoice",
"description": "Matches an invoice to its purchase order and lists line-level differences.",
"tags": ["finance", "invoices"],
"examples": ["Reconcile invoice INV-20931 against its PO."],
"inputModes": ["text/plain", "application/pdf"],
"outputModes": ["application/json"]
},
"requires_scopes": ["recon.run"],
"tenants": "*",
"handler": "recon.v3.reconcile",
"limits": {"max_pdf_pages": 200, "cost_class": "medium"},
"owner": "finance-agents-team"
}Only the skill block ever appears in a card. The rest is server-side: visibility decides which card it appears in, requires_scopes and tenants decide who may use it, handler is the routing target, and state (preview, ga, deprecated or removed) drives the lifecycle described below. Visibility is public, extended or internal. The id is stable across versions of the handler; callers and registries key on it.
Projecting the public and extended cards
Card generation is a pure function of the registry, the caller and the deployment's protocol settings. The public card is served unauthenticated at the well-known path and must contain only what any caller may know. If some skills are only for authenticated partners, set capabilities.extendedAgentCard to true and serve the fuller card from the GetExtendedAgentCard operation. The specification requires that call to fail with UnsupportedOperationError when the flag is false or absent, and defines ExtendedAgentCardNotConfiguredError for an agent that sets the flag but has nothing configured.
def visible(m, caller):
if m["state"] == "removed":
return False
if m["visibility"] == "public":
return True
if m["visibility"] == "extended" and caller is not None:
return (set(m["requires_scopes"]) <= caller.scopes
and (m["tenants"] == "*" or caller.tenant in m["tenants"]))
return False
def build_card(registry, base, caller=None):
card = dict(base) # name, description, version, interfaces, security...
card["skills"] = [m["skill"] | {"id": m["id"]}
for m in registry if visible(m, caller)]
card["capabilities"] = dict(base["capabilities"],
extendedAgentCard=any(m["visibility"] == "extended" for m in registry))
return card
public_card = build_card(REGISTRY, BASE) # served at /.well-known/agent-card.json
partner_card = build_card(REGISTRY, BASE, caller) # returned by GetExtendedAgentCardTwo details matter in practice. First, version the card: bump the card's version whenever the generated skill list changes, so caches that key on it refresh. Second, never let an extended-only skill leak into the public card through a shared description or example; generate, then diff the public card in CI against the previous release and require review for any change.
Routing without a skill id
A detail that surprises people: the A2A SendMessageRequest carries a message, a configuration and metadata, but no field naming the skill the caller wants. Skills are advertisement, not an addressing scheme. The server has to infer which capability a message is for.
Good routers work in layers, cheapest first. Structural signals come first: the media types of the message parts, a structured data part with a recognisable shape, or a task id that continues an existing task already bound to a handler. Then a classifier over the text, restricted to the capabilities this caller is entitled to. Only then a model-based router that reads the candidate skill descriptions. Restricting the candidate set by entitlement before classification is the important step; a router that can pick a capability the caller may not use will eventually pick it.
def route(msg, caller, registry, classify):
if msg.task_id and (bound := TASKS.handler_for(msg.task_id)):
return bound # continuing work stays put
allowed = [m for m in registry
if m["state"] in ("ga", "deprecated", "preview") and visible(m, caller)]
kinds = {part_media_type(p) for p in msg.parts}
fits = [m for m in allowed if kinds <= set(m["skill"]["inputModes"])]
if not fits:
raise A2AError("ContentTypeNotSupportedError")
if len(fits) == 1:
return fits[0]["handler"]
choice, confidence = classify(msg, [m["skill"] for m in fits])
if confidence < 0.6:
return ask_for_clarification(msg, fits) # input-required, not a guess
return next(m["handler"] for m in fits if m["id"] == choice)Some deployments let clients pass a preferred skill id in the request metadata. That can be a useful convention between known partners, but it is not part of the protocol, so treat it as a hint and still check entitlement. When confidence is low, asking the caller to clarify through an input-required task state is cheaper than running the wrong expensive handler. On output, the client's acceptedOutputModes is a SHOULD in the specification: tailor output to it where the capability can, and fail clearly where it cannot.
Advertisement is not authorization
Hiding a skill from a caller's card is not access control. A caller that learned a skill exists from a leaked card, a log or a guess can send a message that a careless router sends straight to it. Every handler must therefore check, at execution time, that the authenticated principal holds the manifest's scopes and tenant, with the same function the card generator used. This is defence in depth, not duplication.
The deeper risk is the confused deputy. An agent acting for a low-privilege caller often holds powerful credentials of its own for its internal tools. If the handler uses those credentials without narrowing them to the caller's rights, the caller gets the agent's privileges by asking politely. Pass the caller's identity down to tool calls, prefer token exchange or downscoped credentials over the agent's service account, and log the principal on every tool invocation. Authentication schemes themselves are covered in A2A authentication.
Extensions: capabilities beyond the core protocol
Extensions let an agent add behaviour the core protocol does not define. The card lists them in capabilities.extensions, each with a URI, a description, a required flag and parameters. The A2A extensions documentation describes four kinds: data-only extensions that add information to the card, profile extensions that add structure or state requirements to the core messages, method extensions that add new RPC methods, and state machine extensions that add task states or transitions.
Activation is per request. A client lists the extension URIs it wants in the A2A-Extensions request header, comma-separated, and the agent replies with the same header naming the extensions it actually activated. If an extension is marked required and the client did not request it, the agent should reject the request; the specification defines ExtensionSupportRequiredError for this case. The URI is the version: the documentation says a breaking change must use a new URI, so put a version segment such as /v1 in it from the start.
def negotiate_extensions(headers, card):
requested = {u.strip() for u in headers.get("A2A-Extensions", "").split(",") if u.strip()}
declared = {e["uri"]: e for e in card["capabilities"].get("extensions", [])}
missing = [u for u, e in declared.items() if e.get("required") and u not in requested]
if missing:
raise A2AError("ExtensionSupportRequiredError", data={"required": missing})
active = sorted(u for u in requested if u in declared) # ignore unknown URIs
return active # echo as the A2A-Extensions response headerEvery required extension shrinks the set of clients that can talk to you, so reserve required for extensions without which a call would be unsafe or meaningless.
Lifecycle: preview, GA, deprecated, removed
Capabilities have a lifecycle, and the manifest's state field should drive it. Preview capabilities appear only in extended cards for opted-in callers. GA capabilities appear wherever visibility allows. Deprecated capabilities stay routable and stay in the card, with the description saying so and naming the replacement, because client agents and registries read descriptions. Removed capabilities disappear from cards but keep a router entry that returns a clear failure naming the replacement, rather than letting messages fall through to the nearest other skill.
Change the meaning of a skill and you have a new skill: give it a new id. Callers key on ids, so silently changing what reconcile-invoice returns breaks them in ways no schema check catches. Protocol version changes are a separate axis, handled per interface; see A2A versioning.
Worked example: a partner-only capability
A finance agent offers invoice lookup publicly and reconciliation only to partners holding the recon.run scope. A new partner's orchestrator fetches the public card and sees only lookup. It authenticates, calls GetExtendedAgentCard, and receives a card that also lists reconcile-invoice, because the partner's token carries the scope. It sends a message with a PDF part and the text reconcile this against the PO.
The router finds no continuing task, filters to capabilities the partner may use, finds two accepting PDFs, and the classifier picks reconcile-invoice with high confidence. The handler re-checks the scope, calls the ERP tool with a token downscoped to the partner's tenant, and returns JSON. Months later the team ships reconcile-invoice-v2 with a different output shape under a new id, marks the old one deprecated with a pointer, and removes it after the registry shows no calls for 30 days. Search and freshness on the registry side are covered in agent registry architecture.
Failure modes
| Failure | Symptom | Prevention |
|---|---|---|
| Hand-edited card | Callers ask for skills that no longer route | Generate cards from the registry; diff in CI |
| Hidden equals secure | A guessed request reaches a partner-only handler | Check scopes and tenant in every handler |
| Confused deputy | Low-privilege callers act with the agent's credentials | Downscope tool credentials to the caller |
| Router picks outside entitlement | Errors or data leaks on ambiguous requests | Filter candidates by entitlement before classifying |
| Meaning changed under a stable id | Downstream agents break without an error | New behaviour, new id; deprecate the old |
| Too many required extensions | Most clients cannot call at all | Require only what safety demands |
Trade-offs
| Decision | Option A | Option B |
|---|---|---|
| Card source | Generated from the registry: cannot drift, needs tooling | Hand-written: quick to start, drifts from the code |
| Routing | Inferred from message content: works with any client, can misroute | Skill-id hint in metadata: precise, only works with partners who adopt the convention |
| Visibility | Public skills: discoverable by registries and new callers | Extended-only: less exposure, callers must authenticate first |
| Extensions | Optional: every client can call | Required: enforces behaviour, excludes clients that lack it |
What to do next
- Inventory what your agent does and separate protocol capabilities, skills and internal tools.
- Write a manifest per skill with state, visibility, scopes, tenants and handler, and generate both cards from it.
- Add a CI check that diffs the generated public card against the last release.
- Build the router to filter by entitlement and media type first, and to ask for clarification when unsure.
- Enforce scopes and tenant in every handler, and pass the caller's identity down to tool calls.
- Version extension URIs, implement A2A-Extensions negotiation, and keep required extensions rare.
- Define the deprecation process: new id for new meaning, deprecated state with a pointer, removal after usage reaches zero.