Finding an A2A agent is only half of discovery. Once a client has an Agent Card, from the well-known path /.well-known/agent-card.json, a registry or configuration, it still has to answer a narrower question: what will this agent actually do for my request, through which interface, with what inputs and outputs, and under which credentials? And it has to keep answering that question as the agent is redeployed, because cards change while workflows that depend on them keep running.

This article covers that step, which we will call capability discovery. Locating cards and choosing between candidates is covered in A2A discovery, and the meaning of each capability flag in A2A capabilities. Here the focus is the client-side machinery: compiling a card into an effective capability profile, folding in the authenticated extended card, checking claims against behaviour, recording what a workflow relies on, and noticing when that stops being true. Field names follow A2A 1.0 as described in the Agent Card specification article.

Advertisement

Three questions behind every delegation

Before a planner hands work to a remote agent, three questions must have answers. Can we talk to it at all: is there an interface whose binding and protocol version we implement, and can we satisfy its security requirements? What will it accept and produce for the specific skill we want: which media types in, which out, which optional operations such as streaming? And will it behave as advertised: a card is a claim written by the provider, sometimes generated from configuration, sometimes stale.

A card answers the first two questions only indirectly, because the answer is spread across several fields with inheritance rules between them. The third question cannot be answered from the card at all. Capability discovery is the process that turns the card into explicit answers, adds evidence, and keeps both current.

From a fetched Agent Card to a capability profile your planner can trustAgent Cardpublic, from discoveryValidateschema, signature, sizeCompileeffective per-skill profileExtended cardafter authenticationif flagProbeverify the claimsCapability lockfilewhat this workflow needsPlanner / routercalls gated by profileA2A callsSendMessage, ...Refreshconditional GET, recompile, diff against lockfileRuntime feedbackUnsupportedOperationError etc. mark the profile staleon scheduleDiscovery answers where an agent is; capability discovery answers what it will do for this request, and keeps that answer current.
The capability discovery pipeline on the client. A validated card is compiled into a per-skill profile, enriched with the extended card where offered, verified with probes and pinned in a lockfile. Refreshes and runtime errors loop back into a diff.

Compiling the effective profile

Compilation resolves the card's indirections once, so the rest of the client never reads raw card fields. The rules come straight from the 1.0 card structure. supportedInterfaces is in the agent's preference order, and a client takes the first entry whose protocolBinding and protocolVersion it supports, copying any declared tenant into requests. A skill's inputModes and outputModes replace defaultInputModes and defaultOutputModes rather than adding to them. A skill can carry its own securityRequirements. Capability flags gate optional operations, and an extension marked required that the client does not understand makes the agent unusable for that client.

SUPPORTED = {("JSONRPC", "1.0"), ("HTTP+JSON", "1.0")}   # what this client implements
KNOWN_EXTENSIONS = {"https://example.com/ext/cost-estimate/v1"}

def compile_profile(card):
    iface = next((i for i in card["supportedInterfaces"]
                  if (i["protocolBinding"], i["protocolVersion"]) in SUPPORTED), None)
    if iface is None:
        return {"usable": False, "reason": "no mutually supported interface"}

    caps = card["capabilities"]
    exts = caps.get("extensions", [])
    missing = [e["uri"] for e in exts if e.get("required") and e["uri"] not in KNOWN_EXTENSIONS]
    if missing:
        return {"usable": False, "reason": f"required extensions not supported: {missing}"}

    skills = {}
    for s in card["skills"]:
        skills[s["id"]] = {
            # A skill's modes REPLACE the card defaults; they are not merged.
            "input": s.get("inputModes") or card["defaultInputModes"],
            "output": s.get("outputModes") or card["defaultOutputModes"],
            "security": s.get("securityRequirements") or card.get("securityRequirements", []),
            "tags": s["tags"],
        }
    return {
        "usable": True,
        "agent_version": card["version"],
        "interface": {"url": iface["url"], "binding": iface["protocolBinding"],
                      "version": iface["protocolVersion"], "tenant": iface.get("tenant")},
        "streaming": bool(caps.get("streaming")),
        "push": bool(caps.get("pushNotifications")),
        "extended_card": bool(caps.get("extendedAgentCard")),
        "extensions": sorted(e["uri"] for e in exts if e["uri"] in KNOWN_EXTENSIONS),
        "skills": skills,
    }

The output is a small, flat structure: one chosen interface, booleans for the optional operations, the extensions both sides know, and a per-skill record of modes, security and tags. Two details are easy to get wrong. Treat a missing optional flag as false, never as unknown-so-try-it. And treat the replacement rule for modes literally: a skill that lists only application/pdf as input does not accept the card's default text/plain, even though the card as a whole does.

Validate before compiling. Reject cards that are not valid JSON for the 1.0 structure, that exceed a size limit you choose, or whose signatures fail verification when you require signed cards. A compiler fed a malformed card produces a confidently wrong profile.

Advertisement

Public card, extended card

The public card is what anyone can read. When capabilities.extendedAgentCard is true, an authenticated caller can request a fuller card with the GetExtendedAgentCard operation, over JSON-RPC by that method name or over the HTTP+JSON binding as GET /extendedAgentCard. Providers use it to show partner-only skills, additional interfaces or different limits to callers who have proved who they are.

Three rules keep this sound. Only ask when the flag is set; otherwise the agent must answer UnsupportedOperationError, and if the flag is set but nothing is configured it answers ExtendedAgentCardNotConfiguredError, which you should treat as fall back to the public card rather than as an outage. Compile the extended card as a whole replacement for the public one for that caller, not as a union, because the provider may intentionally present different security requirements or interfaces. And cache the compiled extended profile per credential identity, not per agent: two tenants of your orchestrator may legitimately see different skills from the same agent, and a shared cache leaks one tenant's view to the other.

Claims are not behaviour: verifying with probes

A card can be wrong without anyone lying. A provider runs the same card template across staging and production while streaming is disabled behind one load balancer, or a skill's output mode was changed in code but not in configuration. The only way to know is to try, carefully.

Probes are small, deliberate calls that check one claim each. They run when a profile is first pinned, after a card change, and on a slow schedule, never on the request path. Useful probes include a streaming call with a trivial request to confirm SendStreamingMessage works when the card says streaming is supported; a request with each media type your workflow will send, to confirm it is accepted rather than rejected with ContentTypeNotSupportedError; and a call without an extension your workflow relies on, to confirm the agent answers with ExtensionSupportRequiredError when that extension is required, which tells you the declaration is enforced.

Probes must be safe. Agree with the provider on a request that has no side effects, or on a dedicated test skill or tenant, and never probe a skill that books, buys or sends something. Respect rate limits; a probe that triggers throttling has measured your client, not the agent. Record each result with a timestamp in the profile as evidence, so the planner can distinguish claimed from verified.

The capability lockfile

A workflow rarely uses everything an agent offers. It depends on one interface, one or two skills, specific media types and perhaps streaming. Write those dependencies down in a lockfile, by analogy with package lockfiles: the subset of the compiled profile this workflow needs, the card version and digest it was verified against, and the probe evidence.

{
  "agent": "https://pricing.partner.example/.well-known/agent-card.json",
  "pinned_at": "2026-10-02T08:30:00Z",
  "card_digest": "sha256:6f1c...e2",
  "agent_version": "3.2.0",
  "interface": {"binding": "JSONRPC", "version": "1.0"},
  "requires": {
    "streaming": true,
    "skills": {
      "itinerary-pricing": {
        "input":  ["application/json"],
        "output": ["application/json"],
        "scopes": ["pricing.quote"]
      }
    },
    "extensions": []
  },
  "verified": {"streaming": "probe ok 2026-10-02", "itinerary-pricing": "probe ok 2026-10-02"}
}

The lockfile changes the question asked on every refresh from 'what does this agent offer now?' to 'does it still offer what we need?' The second question has a clear answer and a clear owner. It also documents the integration for the next engineer: a reviewer can see in one file which remote behaviours a workflow relies on, which is what an A2A threat model needs as input as well.

Refreshing and diffing

Refresh cards on a schedule, using conditional requests where the server provides validators: send the stored ETag in If-None-Match and treat a 304 as no change. A changed version field is a strong hint but not proof either way, so always recompile and compare the compiled profile, not the raw JSON, against the lockfile. Comparing raw JSON flags reordered keys and rewritten descriptions; comparing profiles flags only what affects calls.

def check_against_lock(profile, lock):
    # Classify a recompiled profile against what this workflow depends on.
    if not profile["usable"]:
        return "breaking", [profile["reason"]]
    problems, notes = [], []
    need = lock["requires"]
    if need.get("streaming") and not profile["streaming"]:
        problems.append("streaming no longer advertised")
    for sid, want in need["skills"].items():
        have = profile["skills"].get(sid)
        if have is None:
            problems.append(f"skill {sid} removed (or hidden from this caller)")
            continue
        if not set(want["input"]) <= set(have["input"]):
            problems.append(f"{sid}: input modes narrowed to {have['input']}")
        if not set(want["output"]) & set(have["output"]):
            problems.append(f"{sid}: none of our output modes offered")
    if (profile["interface"]["binding"], profile["interface"]["version"]) != \
       (lock["interface"]["binding"], lock["interface"]["version"]):
        notes.append("preferred interface changed; re-run probes")
    if profile["agent_version"] != lock["agent_version"]:
        notes.append("agent version changed")
    if problems:
        return "breaking", problems
    return ("review", notes) if notes else ("compatible", [])

The classification drives the response. Compatible changes update the cache silently. Review changes trigger probes and notify the owning team. Breaking changes stop new delegations from that workflow and route to an alternative or a human, before a request fails halfway through a multi-agent plan. How often to refresh is the main trade-off: frequent refreshes catch change sooner and cost load on both sides; daily, plus probes only on change, is a reasonable start.

Worked example: a travel planner and a pricing agent

A travel planning orchestrator delegates fare pricing to a partner agent. The lockfile above records its needs: the JSON-RPC 1.0 interface, streaming for long quotes, and the itinerary-pricing skill with JSON in and JSON out under the pricing.quote scope. Over a quarter, the partner changes its card four times.

Card changeProfile effectClassificationAction
Adds skill seat-mapsnew skill entrycompatiblecache update only
Lists HTTP+JSON 1.0 first, JSON-RPC secondchosen interface becomes HTTP+JSON 1.0reviewprobe the new interface, then re-pin
itinerary-pricing outputModes becomes text/markdown onlyoutput no longer includes application/jsonbreakingstop delegating; alert owners; fall back to the secondary pricing agent
Adds a required extension the planner does not knowprofile becomes unusablebreakingsame, plus a ticket to evaluate the extension

Without compilation and a lockfile, the third change surfaces as parse errors in production, possibly after the planner has already reserved seats with another agent. With them, it surfaces as one alert within a refresh interval, and the planner's fallback is a routing decision rather than an incident.

Using the profile at call time

At call time the profile, not the card, decides what to send. Use the chosen interface's URL, send its protocol version in A2A-Version, activate the extensions both sides support through the A2A-Extensions header, copy any tenant value, and pick operations only when their flags are true.

def send(profile, skill_id, message, token, stream_wanted):
    iface = profile["interface"]
    headers = {"A2A-Version": iface["version"], "Authorization": f"Bearer {token}"}
    if profile["extensions"]:
        headers["A2A-Extensions"] = ",".join(profile["extensions"])   # activate what both sides know
    if iface["tenant"]:
        message["tenant"] = iface["tenant"]          # copy verbatim, as declared
    method = "SendStreamingMessage" if (stream_wanted and profile["streaming"]) else "SendMessage"
    try:
        return rpc(iface["url"], method, message, headers)
    except A2AError as e:
        if e.name == "UnsupportedOperationError" and method == "SendStreamingMessage":
            profile["streaming"] = False             # the claim was wrong for this deployment
            mark_stale(profile)                      # refresh and re-diff soon
            return rpc(iface["url"], "SendMessage", message, headers)
        if e.name in ("ContentTypeNotSupportedError", "VersionNotSupportedError"):
            mark_stale(profile)
        raise

Runtime errors are evidence too. An UnsupportedOperationError on a call the profile allowed means the profile is wrong for this deployment: degrade to the fallback operation, mark the profile stale, and let the refresh path recompile and re-diff. A VersionNotSupportedError usually means the agent dropped an interface version you pinned. Feed both into the same diff machinery rather than handling them ad hoc in each workflow.

Treat card content as untrusted input

Every string in a card was written by the provider. Skill names, descriptions and examples are useful to a model that plans delegations, and they are also a prompt-injection channel. Pass them to a planning model as quoted data with length limits, never as instructions, and never let them influence the hard checks in compilation. URLs in a card, such as documentation or icons, should not be fetched automatically by the client. Where your trust model allows it, require signed cards and verify the signature before compiling, so a compromised hosting path cannot silently rewrite capabilities.

Failure modes

  • Merging modes instead of replacing. The client sends text to a skill that accepts only PDFs because it read the card defaults.
  • Shared extended-card cache. One tenant's partner-only skills appear in another tenant's planner.
  • Probes with side effects. A verification call creates real orders or messages. Agree a safe probe with the provider.
  • Raw JSON diffs. Every description edit pages someone, so alerts get muted and the real break is missed.
  • Trusting flags forever. A flag that was true at pin time is false after a redeploy, and nothing notices until a user does.
  • Per-request card fetches. Fetching and compiling on every call adds latency and load; cache compiled profiles and refresh on a schedule.

What to do next

  1. Write a card compiler that picks the interface, applies the mode replacement rule, resolves per-skill security and rejects unknown required extensions.
  2. Validate cards before compiling, with a size limit and signature checks where your trust model requires them.
  3. Fetch the extended card only when the flag is set, compile it as a replacement, and cache it per credential identity.
  4. Agree side-effect-free probes with each provider and record their results as evidence in the profile.
  5. Create a capability lockfile for each workflow that delegates to a remote agent.
  6. Refresh cards with conditional requests, diff compiled profiles against lockfiles, and route breaking changes to a fallback and an owner.
  7. Drive calls from the profile, and feed protocol errors back as stale markers.
  8. Treat every string in a card as untrusted data when it reaches a planning model.
Key takeaway: An Agent Card says where an agent is and what it claims; capability discovery turns that into something a client can depend on. Compile the card once into an explicit per-skill profile, replace it with the extended card for authenticated callers, verify the claims that matter with safe probes, pin what each workflow needs in a lockfile, and diff every refresh and every protocol error against that lockfile. Then a provider's change becomes an alert and a routing decision instead of a failure in the middle of someone's plan.