An A2A agent advertises what it can do in its Agent Card, and client agents decide whether and how to call it from that advertisement. Most guidance stops at the card: which fields exist and how to fill them in. The harder engineering lives behind it. Someone has to decide which skills a given caller may see, which internal handler a request is sent to when the protocol does not say which skill the caller wanted, what happens when a caller tries something it was never shown, and how a capability is retired without breaking the agents that depend on it.

This article covers that server-side capability layer. It separates the three things the word capability is used for, proposes one internal registry from which the public card, the extended card, the request router and the execution policy are all derived, and works through routing, enforcement, extension activation and lifecycle with code. The card's fields are covered field by field in the A2A Agent Card specification; this page assumes them. Statements about the protocol refer to the A2A 1.0 specification as published at a2a-protocol.org.

Advertisement

Three things called capability

Conversations about agent capabilities go wrong because three different layers share one word. Keeping them apart is most of the design.

LayerWhere it appearsWho consumes itChanges when
Protocol capabilitiescapabilities in the card: streaming, push notifications, extended card, extensionsClient SDKs, before choosing an operationYou deploy new protocol features
Skillsskills in the card: id, name, description, tags, examples, media typesRouting models, registries, human integratorsYou add, change or retire a business function
Internal capabilitiesNowhere public: tools, model endpoints, data sources and the permissions to use themYour own runtime and policy engineEvery release

Protocol capabilities are switches: each flag gates an operation, and a client must check it before trying. Skills are descriptions written for routing: they tell a caller what to ask for, not how the agent does it. Internal capabilities are implementation detail and should never leak into the card. One skill such as reconcile an invoice may use five internal tools; one tool may serve several skills.

One registry, four projections

The failure pattern to avoid is the hand-written card: a JSON file edited by whoever shipped the last feature, drifting away from what the code actually supports. A skill that is advertised but not routable produces confused callers; a skill that is routable but not advertised is an undocumented API that will be found and depended on.

Instead, keep one internal registry of capability manifests and generate everything from it. The public card is the projection of manifests marked public. The extended card is the projection a specific authenticated caller is entitled to. The router's index is the set of handlers that caller may reach. The execution policy is the set of scopes and limits each handler enforces. Because all four read the same records, they cannot disagree.

One capability registry, four projectionsCapability registrymanifests: skill, scopes, modes, statePublic Agent Cardvisibility = publicExtended Agent Cardper authenticated callerRouter indexentitled capabilities onlyExecution policyscopes, tenant, budgetClient agentreads card, sends messageA2A endpointversion, extensions, authCapability handlertools, models, dataSendMessagerouted + checkeddiscoveryCards advertise, the router chooses, the policy layer decides. All three read the same registry.
The capability registry is the source of truth. The public card, the extended card, the router and the execution policy are all generated from it, so advertisement and enforcement cannot drift apart.
Advertisement

The capability manifest

A manifest records the skill as it should be advertised plus everything the server needs to route and enforce it. A minimal one looks like this:

{
  "id": "reconcile-invoice",
  "state": "ga",
  "visibility": "extended",
  "skill": {
    "name": "Reconcile an invoice",
    "description": "Matches an invoice to its purchase order and lists line-level differences.",
    "tags": ["finance", "invoices"],
    "examples": ["Reconcile invoice INV-20931 against its PO."],
    "inputModes": ["text/plain", "application/pdf"],
    "outputModes": ["application/json"]
  },
  "requires_scopes": ["recon.run"],
  "tenants": "*",
  "handler": "recon.v3.reconcile",
  "limits": {"max_pdf_pages": 200, "cost_class": "medium"},
  "owner": "finance-agents-team"
}

Only the skill block ever appears in a card. The rest is server-side: visibility decides which card it appears in, requires_scopes and tenants decide who may use it, handler is the routing target, and state (preview, ga, deprecated or removed) drives the lifecycle described below. Visibility is public, extended or internal. The id is stable across versions of the handler; callers and registries key on it.

Projecting the public and extended cards

Card generation is a pure function of the registry, the caller and the deployment's protocol settings. The public card is served unauthenticated at the well-known path and must contain only what any caller may know. If some skills are only for authenticated partners, set capabilities.extendedAgentCard to true and serve the fuller card from the GetExtendedAgentCard operation. The specification requires that call to fail with UnsupportedOperationError when the flag is false or absent, and defines ExtendedAgentCardNotConfiguredError for an agent that sets the flag but has nothing configured.

def visible(m, caller):
    if m["state"] == "removed":
        return False
    if m["visibility"] == "public":
        return True
    if m["visibility"] == "extended" and caller is not None:
        return (set(m["requires_scopes"]) <= caller.scopes
                and (m["tenants"] == "*" or caller.tenant in m["tenants"]))
    return False

def build_card(registry, base, caller=None):
    card = dict(base)                    # name, description, version, interfaces, security...
    card["skills"] = [m["skill"] | {"id": m["id"]}
                      for m in registry if visible(m, caller)]
    card["capabilities"] = dict(base["capabilities"],
                                extendedAgentCard=any(m["visibility"] == "extended" for m in registry))
    return card

public_card = build_card(REGISTRY, BASE)              # served at /.well-known/agent-card.json
partner_card = build_card(REGISTRY, BASE, caller)     # returned by GetExtendedAgentCard

Two details matter in practice. First, version the card: bump the card's version whenever the generated skill list changes, so caches that key on it refresh. Second, never let an extended-only skill leak into the public card through a shared description or example; generate, then diff the public card in CI against the previous release and require review for any change.

Routing without a skill id

A detail that surprises people: the A2A SendMessageRequest carries a message, a configuration and metadata, but no field naming the skill the caller wants. Skills are advertisement, not an addressing scheme. The server has to infer which capability a message is for.

Good routers work in layers, cheapest first. Structural signals come first: the media types of the message parts, a structured data part with a recognisable shape, or a task id that continues an existing task already bound to a handler. Then a classifier over the text, restricted to the capabilities this caller is entitled to. Only then a model-based router that reads the candidate skill descriptions. Restricting the candidate set by entitlement before classification is the important step; a router that can pick a capability the caller may not use will eventually pick it.

def route(msg, caller, registry, classify):
    if msg.task_id and (bound := TASKS.handler_for(msg.task_id)):
        return bound                                     # continuing work stays put
    allowed = [m for m in registry
               if m["state"] in ("ga", "deprecated", "preview") and visible(m, caller)]
    kinds = {part_media_type(p) for p in msg.parts}
    fits = [m for m in allowed if kinds <= set(m["skill"]["inputModes"])]
    if not fits:
        raise A2AError("ContentTypeNotSupportedError")
    if len(fits) == 1:
        return fits[0]["handler"]
    choice, confidence = classify(msg, [m["skill"] for m in fits])
    if confidence < 0.6:
        return ask_for_clarification(msg, fits)          # input-required, not a guess
    return next(m["handler"] for m in fits if m["id"] == choice)

Some deployments let clients pass a preferred skill id in the request metadata. That can be a useful convention between known partners, but it is not part of the protocol, so treat it as a hint and still check entitlement. When confidence is low, asking the caller to clarify through an input-required task state is cheaper than running the wrong expensive handler. On output, the client's acceptedOutputModes is a SHOULD in the specification: tailor output to it where the capability can, and fail clearly where it cannot.

Advertisement is not authorization

Hiding a skill from a caller's card is not access control. A caller that learned a skill exists from a leaked card, a log or a guess can send a message that a careless router sends straight to it. Every handler must therefore check, at execution time, that the authenticated principal holds the manifest's scopes and tenant, with the same function the card generator used. This is defence in depth, not duplication.

The deeper risk is the confused deputy. An agent acting for a low-privilege caller often holds powerful credentials of its own for its internal tools. If the handler uses those credentials without narrowing them to the caller's rights, the caller gets the agent's privileges by asking politely. Pass the caller's identity down to tool calls, prefer token exchange or downscoped credentials over the agent's service account, and log the principal on every tool invocation. Authentication schemes themselves are covered in A2A authentication.

Extensions: capabilities beyond the core protocol

Extensions let an agent add behaviour the core protocol does not define. The card lists them in capabilities.extensions, each with a URI, a description, a required flag and parameters. The A2A extensions documentation describes four kinds: data-only extensions that add information to the card, profile extensions that add structure or state requirements to the core messages, method extensions that add new RPC methods, and state machine extensions that add task states or transitions.

Activation is per request. A client lists the extension URIs it wants in the A2A-Extensions request header, comma-separated, and the agent replies with the same header naming the extensions it actually activated. If an extension is marked required and the client did not request it, the agent should reject the request; the specification defines ExtensionSupportRequiredError for this case. The URI is the version: the documentation says a breaking change must use a new URI, so put a version segment such as /v1 in it from the start.

def negotiate_extensions(headers, card):
    requested = {u.strip() for u in headers.get("A2A-Extensions", "").split(",") if u.strip()}
    declared = {e["uri"]: e for e in card["capabilities"].get("extensions", [])}
    missing = [u for u, e in declared.items() if e.get("required") and u not in requested]
    if missing:
        raise A2AError("ExtensionSupportRequiredError", data={"required": missing})
    active = sorted(u for u in requested if u in declared)   # ignore unknown URIs
    return active          # echo as the A2A-Extensions response header

Every required extension shrinks the set of clients that can talk to you, so reserve required for extensions without which a call would be unsafe or meaningless.

Lifecycle: preview, GA, deprecated, removed

Capabilities have a lifecycle, and the manifest's state field should drive it. Preview capabilities appear only in extended cards for opted-in callers. GA capabilities appear wherever visibility allows. Deprecated capabilities stay routable and stay in the card, with the description saying so and naming the replacement, because client agents and registries read descriptions. Removed capabilities disappear from cards but keep a router entry that returns a clear failure naming the replacement, rather than letting messages fall through to the nearest other skill.

Change the meaning of a skill and you have a new skill: give it a new id. Callers key on ids, so silently changing what reconcile-invoice returns breaks them in ways no schema check catches. Protocol version changes are a separate axis, handled per interface; see A2A versioning.

Worked example: a partner-only capability

A finance agent offers invoice lookup publicly and reconciliation only to partners holding the recon.run scope. A new partner's orchestrator fetches the public card and sees only lookup. It authenticates, calls GetExtendedAgentCard, and receives a card that also lists reconcile-invoice, because the partner's token carries the scope. It sends a message with a PDF part and the text reconcile this against the PO.

The router finds no continuing task, filters to capabilities the partner may use, finds two accepting PDFs, and the classifier picks reconcile-invoice with high confidence. The handler re-checks the scope, calls the ERP tool with a token downscoped to the partner's tenant, and returns JSON. Months later the team ships reconcile-invoice-v2 with a different output shape under a new id, marks the old one deprecated with a pointer, and removes it after the registry shows no calls for 30 days. Search and freshness on the registry side are covered in agent registry architecture.

Failure modes

FailureSymptomPrevention
Hand-edited cardCallers ask for skills that no longer routeGenerate cards from the registry; diff in CI
Hidden equals secureA guessed request reaches a partner-only handlerCheck scopes and tenant in every handler
Confused deputyLow-privilege callers act with the agent's credentialsDownscope tool credentials to the caller
Router picks outside entitlementErrors or data leaks on ambiguous requestsFilter candidates by entitlement before classifying
Meaning changed under a stable idDownstream agents break without an errorNew behaviour, new id; deprecate the old
Too many required extensionsMost clients cannot call at allRequire only what safety demands

Trade-offs

DecisionOption AOption B
Card sourceGenerated from the registry: cannot drift, needs toolingHand-written: quick to start, drifts from the code
RoutingInferred from message content: works with any client, can misrouteSkill-id hint in metadata: precise, only works with partners who adopt the convention
VisibilityPublic skills: discoverable by registries and new callersExtended-only: less exposure, callers must authenticate first
ExtensionsOptional: every client can callRequired: enforces behaviour, excludes clients that lack it

What to do next

  1. Inventory what your agent does and separate protocol capabilities, skills and internal tools.
  2. Write a manifest per skill with state, visibility, scopes, tenants and handler, and generate both cards from it.
  3. Add a CI check that diffs the generated public card against the last release.
  4. Build the router to filter by entitlement and media type first, and to ask for clarification when unsure.
  5. Enforce scopes and tenant in every handler, and pass the caller's identity down to tool calls.
  6. Version extension URIs, implement A2A-Extensions negotiation, and keep required extensions rare.
  7. Define the deprecation process: new id for new meaning, deprecated state with a pointer, removal after usage reaches zero.
Key takeaway: An A2A card is an advertisement, and the protocol leaves routing and enforcement to you. Keep one capability registry as the source of truth, generate the public card, the extended card, the router's candidate set and the execution policy from it, route by structure and entitlement before any model guesses, enforce scopes in every handler with the caller's identity carried through to tools, negotiate extensions per request, and retire capabilities by id with a clear deprecation path.