Once an organisation runs more than a handful of agents, the question stops being how to call an agent and becomes which agent to call. An agent registry answers it. It is a catalogue of Agent Cards that other agents and people can search by skill, tag or input type, filtered by what the caller is allowed to see, and annotated with whether each agent is actually up.

This article treats the registry as a production service. It starts with what A2A defines and leaves open, then covers architecture, data model, verification, search, liveness, policy, federation, sizing and failure modes. The card format itself, and how a client caches and trusts a single card, are covered in the Agent Card deep dive.

Advertisement

What A2A standardises, and what it leaves to you

A2A gives every agent a self-description, the Agent Card, and a standard place to publish it: https://{domain}/.well-known/agent-card.json, following RFC 8615 for well-known URIs. The card lists the agent's name, description, version, provider, the interfaces it can be reached on (supportedInterfaces, each with a URL and protocol binding), its capabilities such as streaming and push notifications, the security schemes it accepts, default input and output modes, and a list of skills. Each skill has an id, name, description, tags, examples and its own input and output modes. A card can also carry signatures: JWS signatures (RFC 7515) computed over a canonical JSON form of the card (JCS, RFC 8785).

What A2A does not define is a registry. The discovery documentation describes curated registries as one discovery strategy and states plainly that the current specification does not prescribe a standard API for them. So everything in this article about the registry's own API, storage and states is design guidance, not protocol. If other organisations' clients will use your registry, document its API yourself and expect future conventions to change it.

The reference architecture

The design has two narrow waists. Every write, whether pushed by an owner or pulled by a crawler, goes through one verifier. Every read goes through one policy filter. Between them sit an append-only card store, a search index derived from it, and a health prober that annotates entries. The registry is on the discovery path only: once a client has picked an agent it calls that agent directly, and the agent enforces its own authentication. A registry that proxies task traffic becomes a single point of failure.

An agent registry: cards come in through one verified pipeline, queries go out through one policy filterAgent ownerregister / updateCrawlerpolls well-known URLsVerifierschema, domain, JWSCard storeappend-only versionsSearch indexskills, tags, vectorsHealth proberreachability, latencyQuery APIfind, get, watchPolicy filtercaller entitlementsClient agentcaches resultsAudit logwho saw and changed whatacceptindexstatusresultsAfter discovery the client calls the agent directly; the registry never proxies task traffic.
Reference architecture. The card store is the source of truth; the index and health data are derived and can be rebuilt.

Keep the registry's own state small and derivable. The card store holds exactly what each agent published, with a revision number, hash, ETag and fetch time. Everything else, including skill rows, embeddings and ranking features, can be rebuilt from it. That makes index corruption or a change of embedding model a rebuild, not a migration.

Advertisement

The data model

An agent is identified by its card URL, not by the name inside the card. Names collide; the URL is tied to a domain you can check. Store every accepted card as a new immutable revision, so you can answer what an agent advertised last Tuesday and roll back a bad publish. Denormalise skills into their own rows so search does not parse JSON at query time.

-- Illustrative schema (PostgreSQL). A2A 1.0 does not define a registry API or storage format.
CREATE TABLE agent (
  agent_id      uuid PRIMARY KEY,
  tenant_id     text NOT NULL,
  card_url      text NOT NULL UNIQUE,      -- https://host/.well-known/agent-card.json
  owner_team    text NOT NULL,
  state         text NOT NULL DEFAULT 'pending',  -- pending|active|degraded|stale|delisted
  current_rev   bigint,
  created_at    timestamptz NOT NULL DEFAULT now()
);

CREATE TABLE card_revision (
  agent_id      uuid REFERENCES agent,
  rev           bigint,
  card_json     jsonb NOT NULL,            -- exactly what the agent published
  card_sha256   bytea NOT NULL,
  etag          text,
  signed_by     text,                      -- key id if a signature verified, else NULL
  fetched_at    timestamptz NOT NULL,
  PRIMARY KEY (agent_id, rev)
);

CREATE TABLE skill (                       -- denormalised for search
  agent_id      uuid,
  skill_id      text,
  name          text,
  description   text,
  tags          text[],
  input_modes   text[],
  output_modes  text[],
  embedding     vector(768),               -- optional, pgvector
  PRIMARY KEY (agent_id, skill_id)
);

The state column is the registry's own lifecycle, not an A2A field. A new entry is pending until the first health probe passes, active when healthy, degraded when probes are slow or failing intermittently, stale when the card could not be refetched for a set period, and delisted when an owner or administrator removes it.

Ingest and verification

Offer two ways in. Push registration lets an owner submit a card URL through an authenticated call, which suits CI pipelines that register a new version on deploy. Pull crawling refetches known URLs on a schedule, which catches cards that changed without anyone telling you. Both paths submit the URL, never the card body: the registry fetches the card itself from the well-known location, so a registrant cannot claim a card for a domain it does not control.

Verification then runs in a fixed order. Fetch over HTTPS only, with a timeout and a byte limit. Validate the JSON against the card schema, including required fields, types and reasonable size limits on descriptions and skill lists. Check that every interface URL is on a host the registrant is allowed to advertise, usually the card's own domain or an allow-list for the tenant; otherwise an agent can point callers at someone else's endpoint. Reject duplicate skill IDs. If the card is signed, verify the signatures against keys you trust for that tenant; if tenant policy requires signatures, reject unsigned cards.

MAX_CARD_BYTES = 256 * 1024

def ingest(card_url, tenant, submitted_by):
    assert card_url.startswith("https://")
    resp = http_get(card_url, timeout=5, max_bytes=MAX_CARD_BYTES,
                    headers={"If-None-Match": stored_etag(card_url)})
    if resp.status == 304:
        return touch(card_url)                # unchanged
    card = json.loads(resp.body)

    errors = validate_against_schema(card)   # required fields, types, sizes
    for iface in card.get("supportedInterfaces", []):
        if host_of(iface["url"]) not in allowed_hosts_for(card_url, tenant):
            errors.append("interface URL outside the card's own domain")

    key_id = None
    if card.get("signatures"):
        key_id = verify_jws_signatures(card, trusted_keys(tenant))  # JWS over JCS form
        if key_id is None:
            errors.append("signature present but not verifiable")
    elif policy(tenant).require_signed_cards:
        errors.append("unsigned card rejected by tenant policy")

    if errors:
        audit("ingest_rejected", card_url, submitted_by, errors)
        return "rejected", errors

    with transaction():
        rev = append_revision(card_url, card, sha256(resp.body), resp.etag, key_id)
        reindex_skills(card_url, card["skills"])   # delete + insert, same transaction
        set_state(card_url, "pending_probe")
    audit("ingest_accepted", card_url, submitted_by, rev)
    return "accepted", rev

Index updates share the revision's transaction, so search never mixes revisions. Audit rejections with reasons; a card that suddenly fails validation often signals a broken deploy.

Search: matching a need to a skill

Callers know what they need, not an agent's name. So search works over skills, in two stages. The first stage applies hard filters that must hold: tenant visibility, entry state, required tags, and required input or output modes such as application/pdf. The second stage ranks what is left, using a mix of lexical scoring over skill names and descriptions and vector similarity over embeddings of the description and examples.

def find_agents(caller, skill_text=None, tags=(), input_mode=None, limit=10):
    q = Query().where(state__in=("active", "degraded"))
    q = q.where(tenant_id__in=caller.visible_tenants)          # hard filter, never ranked
    if tags:
        q = q.where(skill_tags__contains_all=list(tags))
    if input_mode:
        q = q.where(skill_input_modes__contains=input_mode)
    candidates = q.limit(200).fetch()

    if skill_text:
        v = embed(skill_text)
        for c in candidates:
            c.score = 0.7 * cosine(v, c.skill_embedding) + 0.3 * bm25(skill_text, c.skill_description)
    for c in candidates:
        if c.state == "degraded":
            c.score -= 0.2                                     # demote, do not hide
    visible = [c for c in candidates if entitled(caller, c)]   # policy decision point
    audit("search", caller.id, len(visible))
    return sorted(visible, key=lambda c: -c.score)[:limit]

Two rules keep search trustworthy. First, never let ranking do the work of policy: an entry the caller may not see is filtered out before scoring, not merely ranked low. Second, return enough for the caller to decide, including the card URL, the matched skill, the card revision and the health state, rather than a single answer. Several candidates let an orchestrator fail over.

Liveness and freshness

A card says what an agent can do; it does not say whether the agent is running. The health prober fills that gap by periodically making a cheap request to each active agent, such as a conditional fetch of its card, and recording success, latency and consecutive failures. It measures reachability from the registry's network, not every caller's, so use it to demote, not to route.

A workable policy is: probe every 30 to 60 seconds; mark an entry degraded after three consecutive failures or when the latency percentile crosses a threshold; mark it stale when the card has not been refetched successfully for a day; and never delist automatically. Degraded entries still appear in results but lower down and flagged; a registry that hides agents during its own network trouble causes an outage.

Policy, entitlements and audit

Discovery is itself sensitive. Knowing that a payroll agent exists, and what skills it has, is information an attacker would like. The policy filter decides, for each caller identity, which entries it may see, based on tenant, team, data classification tags or explicit grants. Authenticate every query; anonymous search should see only entries explicitly marked public.

A2A also provides an authenticated extended card: an agent can advertise capabilities.extendedAgentCard and return a more detailed card through the GetExtendedAgentCard operation to authenticated clients. A registry should index only the public card unless it holds credentials that entitle it to more, and even then should apply the same visibility rules to the extra detail, or it becomes a way to read cards without authorisation.

Visibility in the registry is not permission to call. The agent must still authenticate and authorise every request using the schemes in its card; see A2A authentication. Audit both sides: who registered, changed or delisted each entry, and who searched for what. Keep audit records outside the registry database.

Tenancy and federation

In a shared platform, make tenant a column on every row and a filter on every query, enforced in the data access layer rather than in each handler. The patterns match those in A2A multi-tenancy.

Across organisations, a single global registry is rarely acceptable. Federation is the usual answer: each organisation runs its own registry and exports a subset of entries to partners, either by letting partners query a scoped endpoint or by publishing a signed export that partners import. Imported entries keep their origin and are re-verified locally, because a partner's registry vouching for a card is weaker than fetching and checking the card yourself. Because the protocol does not define a registry API, federation is a bilateral agreement today; document the format and version it.

Worked example: sizing a registry for 2,000 agents

Take an enterprise with 2,000 registered agents averaging 6 skills each, so 12,000 skill rows. Cards average 8 KB, so the current revisions are about 16 MB; with a year of history at one change per agent per week, revisions add about 800 MB. This fits comfortably in one PostgreSQL instance with a replica. Embeddings at 768 float32 dimensions are about 3 KB per skill, 36 MB in total.

Probing each agent every 30 seconds is about 67 requests per second from the prober, trivial for the registry and negligible for agents if the probe is a conditional card fetch answered with 304.

Query load is dominated by orchestrators. If 500 orchestrator instances each search on every new task and handle 20 tasks per second, that is 10,000 queries per second, which is too many for vector search on every call. Clients should cache results for the same query for 30 to 60 seconds and refresh in the background; that typically cuts load by one to two orders of magnitude. If the registry is down, clients should keep using cached results; test for that.

Failure modes

SymptomCauseFix
Callers routed to an attacker's endpointInterface URL on a foreign host acceptedEnforce host allow-lists; fetch cards yourself; verify signatures
Two agents with the same nameName used as identityIdentify by card URL; show provider and domain
Search returns a dead agentNo health data, or probes ignored in rankingProbe, demote degraded, return several candidates
Whole fleet stops when registry is downClients query on every call with no cacheClient-side caching with stale-while-error
Private agents visible to other teamsVisibility applied after ranking, or not at allFilter by entitlement before scoring
Index shows skills the card no longer hasIndex updated outside the revision transactionUpdate both in one transaction; rebuild on drift
Mass delisting during a network blipAutomatic delist on probe failureDemote only; delist by human action

Keep stored revisions verbatim and version your parser, as in A2A versioning. When a chosen agent fails, fall back to the next candidate behind a breaker, as in the A2A circuit breaker article.

What to do next

  1. Write down, for your organisation, that A2A defines cards and the well-known path but no registry API, and publish your registry's API and states as your own versioned contract.
  2. Identify agents by card URL and store every accepted card as an immutable revision with hash, ETag and fetch time.
  3. Build one verifier: HTTPS fetch with limits, schema validation, interface host checks, duplicate skill checks and signature verification where policy requires it.
  4. Split search into hard filters (tenant, state, tags, modes) and ranking, and apply entitlements before ranking.
  5. Add a health prober that demotes but never delists automatically, and return several candidates with health annotations.
  6. Make clients cache results and keep working when the registry is unavailable; test that with the registry switched off.
  7. Audit registrations, changes and searches to a store outside the registry.
Key takeaway: A2A standardises the Agent Card and where to publish it, but not the registry, so the registry is your own versioned service. Identify agents by card URL, fetch and verify cards yourself, store immutable revisions and derive the search index from them. Search skills with hard filters first and ranking second, apply entitlements before ranking, annotate results with health rather than hiding them, keep the registry off the task path, and make sure clients keep working from cache when it is down.