A multi-agent system in ADK Java is usually a tree fixed at build time. A root LlmAgent lists its sub-agents, and the model transfers control to one of them by name. That works while every agent is compiled into the same service. It stops working once other teams run agents as separate services that come and go, change their skills, or move endpoints. You then need dynamic discovery: finding out at run time which agents exist, what they can do and where they live, and getting that information into your agent tree without restarting.

This article covers the discovery layer itself. You will learn where agent descriptions come from, how to cache and refresh Agent2Agent (A2A) agent cards correctly, how to match a task to capabilities, and how to rebuild ADK agent trees from versioned snapshots so live sessions are never disturbed. Two neighbouring topics have their own articles. Building the A2A client and RemoteA2AAgent is shown in ADK Java and A2A. Deciding which third-party agents may be admitted at all is covered in agent marketplace patterns. The protocol-level view of discovery, independent of ADK, is in A2A agent discovery.

What an agent publishes and what ADK consumes

The A2A protocol describes each agent with an agent card, a JSON document listing the agent's name, description, service URL, version, capabilities such as streaming, security schemes and default modes, plus its skills: per-skill ids, names, descriptions, tags, examples and optional mode overrides. Discovery has three recognised routes. The card can sit at the RFC 8615 path /.well-known/agent-card.json on the agent's own host. Clients can query a curated registry. Or the client can be configured directly with the details. The protocol does not tell you which to use. Production systems usually combine all three.

The ADK Java side has two pieces. A2ACardResolver fetches and parses a card. RemoteA2AAgent wraps a card and an A2A client so that the remote agent behaves like a local BaseAgent in the tree. Remote turns come back as ADK events in your session. The Java SDK implements A2A protocol 0.3 at the time of writing. Note one property of the builder: if you do not pass a card, it asks the client for one during construction. That is a network call at build time, which the design below avoids.

Architecture: a control loop and a snapshot

The central rule is that discovery runs in the background and the request path reads an immutable snapshot. Nothing in a user's turn should wait on a card fetch, and nothing in a card fetch should be able to change a turn that has already started.

Discovery as a background control loop; the request path only reads an immutable snapshotStatic configknown base URLsDomain probing/.well-known/agent-card.jsonRegistry queryyour catalog APIPlatform labelsservice mesh, k8sResolvermerge, dedupe, allowlistCard cacheconditional GET, TTLMatcherskills, modes, policySnapshotversion + digestendpointscardsselectedTree factoryone root per digestrequest pathNew sessionpin current digestRoot LlmAgentbuilt once per digestRemoteA2AAgentA2A call to the peerCards change on the server's schedule; sessions change trees only when they start.
Sources feed a resolver; the card cache and matcher produce a versioned snapshot; the tree factory builds one root agent per snapshot digest, and each new session pins the digest current at its start.
  • Sources each produce endpoint records: a name, a base URL and which source produced it. Static configuration covers the agents you depend on. Domain probing checks a list of partner domains. A registry query covers an internal catalog. Platform labels, for example Kubernetes services carrying a label your team defines, cover agents that scale and move.
  • The resolver merges sources, removes duplicates by base URL, and applies an allowlist of hosts. A source can propose an endpoint, but only the allowlist can admit it.
  • The card cache fetches cards, honours HTTP caching, and remembers failures.
  • The matcher selects agents whose skills fit your application, and applies data and tenant policy.
  • The snapshot is an immutable list of selected agents and their cards, with a version number and a digest computed from card hashes.
  • The tree factory builds a root agent for each digest, once, and caches it.

A card cache that respects HTTP

On caching, the A2A guidance is ordinary HTTP: the publisher states a lifetime in Cache-Control and tags each version of the card with an ETag; the consumer keeps the tag and revalidates with If-None-Match instead of downloading again. The resolver call shown in the A2A quickstart returns only a parsed card, without headers, so the cache below makes its own conditional request with java.net.http to detect change. It re-parses through the resolver only when the content actually changed:

public final class CardCache {
  public record Entry(AgentCard card, String etag, String sha256,
                      Instant expiresAt, int failures) {}

  private final HttpClient http = HttpClient.newBuilder()
      .connectTimeout(Duration.ofSeconds(3)).build();
  private final ConcurrentHashMap<URI, Entry> entries = new ConcurrentHashMap<>();
  private final ConcurrentHashMap<URI, CompletableFuture<Entry>> inflight = new ConcurrentHashMap<>();

  /** Single-flight refresh: concurrent callers for one URL share one fetch. */
  public CompletableFuture<Entry> refresh(URI base) {
    CompletableFuture<Entry> mine = new CompletableFuture<>();
    CompletableFuture<Entry> running = inflight.putIfAbsent(base, mine);
    if (running != null) return running;              // someone is already fetching
    CompletableFuture.supplyAsync(() -> fetch(base)).whenComplete((e, err) -> {
      inflight.remove(base, mine);                     // remove before completing
      if (err != null) mine.completeExceptionally(err); else mine.complete(e);
    });
    return mine;
  }

  private Entry fetch(URI base) {
    Entry old = entries.get(base);
    URI cardUrl = base.resolve("/.well-known/agent-card.json");
    HttpRequest.Builder req = HttpRequest.newBuilder(cardUrl).timeout(Duration.ofSeconds(5)).GET();
    if (old != null && old.etag() != null) req.header("If-None-Match", old.etag());
    try {
      HttpResponse<byte[]> res = http.send(req.build(), HttpResponse.BodyHandlers.ofByteArray());
      Instant expires = Instant.now().plus(maxAge(res).orElse(Duration.ofMinutes(5)));
      if (res.statusCode() == 304 && old != null) {
        return put(base, new Entry(old.card(), old.etag(), old.sha256(), expires, 0));
      }
      if (res.statusCode() != 200) throw new IOException("card HTTP " + res.statusCode());
      String hash = sha256Hex(res.body());
      AgentCard card = (old != null && hash.equals(old.sha256()))
          ? old.card()                                   // same bytes, keep the parsed card
          : new A2ACardResolver(new JdkA2AHttpClient(), base.toString(), cardUrl.toString())
              .getAgentCard();                           // changed: parse via the SDK
      String etag = res.headers().firstValue("ETag").orElse(null);
      return put(base, new Entry(card, etag, hash, expires, 0));
    } catch (Exception e) {
      if (old == null) throw new CompletionException(e);   // never seen: surface it
      Duration backoff = Duration.ofSeconds(Math.min(300, 5L << Math.min(old.failures(), 6)));
      return put(base, new Entry(old.card(), old.etag(), old.sha256(),
          Instant.now().plus(backoff), old.failures() + 1)); // serve stale, retry later
    }
  }

  private Entry put(URI base, Entry e) { entries.put(base, e); return e; }
  // maxAge(...) parses Cache-Control; sha256Hex(...) is a standard MessageDigest helper.
}

Four behaviours are deliberate. Refreshes are single-flight, so a hundred threads noticing an expired entry cause one request. On error the cache serves stale and backs off exponentially, because a peer whose card endpoint is briefly down is usually still serving tasks. A failure counter is kept so that health logic can decide when stale becomes unusable. And the change test is the byte hash, not the ETag, because ETags are optional and some servers emit a new one on every response. One race remains. The card can change between the conditional GET and the parse, so the stored hash may describe slightly older bytes. The next cycle detects and corrects that, which is acceptable for a background loop.

Matching capabilities, not names

A discovered agent is not automatically a useful one. The matcher turns cards into candidates, using hard constraints only. Ranking agents per request belongs to the router, and the query router article covers that layer.

public record Need(String tag, Set<String> inputModes, Set<String> outputModes) {}

static boolean satisfies(AgentCard card, Need need) {
  return card.skills().stream().anyMatch(skill ->
      skill.tags() != null && skill.tags().contains(need.tag())
      && modes(skill.inputModes(), card.defaultInputModes()).containsAll(need.inputModes())
      && modes(skill.outputModes(), card.defaultOutputModes()).containsAll(need.outputModes()));
}

static Set<String> modes(List<String> skillModes, List<String> defaults) {
  return new HashSet<>(skillModes == null || skillModes.isEmpty() ? defaults : skillModes);
}

The fallback from skill modes to card defaults follows the card's structure: a skill that lists no modes inherits the agent's defaults. Treat tags as a vocabulary your organisation owns, published and reviewed, rather than free text. Two teams that both tag a skill refund should mean the same thing. Never let a card's description flow unreviewed into the router's instruction. Card text is written by someone else and is an injection surface. Use a description from your own configuration, as the agent registry design article recommends.

Snapshots, digests and session pinning

The snapshot turns a moving set of cards into stable, comparable values. Sort the selected agents by your local name, hash the list of name and card-hash pairs, and use that digest as the identity of an agent tree. The factory builds one root per digest and caches it. Each session records the digest current when it started:

public record Selected(String localName, String reviewedDescription, Entry entry) {}
public record Snapshot(long version, String digest, List<Selected> agents) {}

public final class TreeFactory {
  private final Map<String, LlmAgent> roots = new ConcurrentHashMap<>();
  private final Map<String, String> sessionDigest = new ConcurrentHashMap<>();
  private final AtomicReference<Snapshot> current = new AtomicReference<>();
  private final Function<AgentCard, Client> clients;   // built as in the A2A integration article

  public void publish(Snapshot s) { current.set(s); }   // called by the discovery loop

  public LlmAgent rootFor(String sessionId) {
    String digest = sessionDigest.computeIfAbsent(sessionId, id -> current.get().digest());
    return roots.computeIfAbsent(digest, d -> build(snapshotFor(d)));
  }

  private LlmAgent build(Snapshot s) {
    List<BaseAgent> remotes = s.agents().stream()
        .map(a -> (BaseAgent) RemoteA2AAgent.builder()
            .name(a.localName())                       // our identifier, never card.name()
            .description(a.reviewedDescription())      // reviewed text, not the card's
            .agentCard(a.entry().card())               // cached: no fetch during build
            .a2aClient(clients.apply(a.entry().card()))
            .build())
        .toList();
    return LlmAgent.builder()
        .name("root_agent")
        .model("gemini-2.5-flash")
        .instruction("Route each request to the one specialist whose description fits it.")
        .subAgents(remotes)
        .build();
  }
  // snapshotFor(d) looks up retained snapshots; evict roots no live session pins.
}

Session pinning matters because a session's history refers to agents by name. If a new snapshot drops refunds_agent halfway through a conversation, the model may try to transfer to a name that no longer exists. With pinning, the old tree keeps serving that session until it ends, and new sessions get the new tree. Retain old snapshots until no live session pins them, then evict. A user who wants a newly discovered agent starts a new session, or your application offers an explicit upgrade at a safe turn boundary. Hand the root to your Runner as usual; in tests, an InMemoryRunner is enough.

Worked example: a morning of changes

Follow one change through the system. A support orchestrator depends on a billing agent from static configuration and discovers specialists from platform labels. At 10:00 the snapshot has digest A with billing_agent and returns_agent. At 10:05 a team deploys a warranty agent with the label. The resolver sees a new endpoint, and the host is on the allowlist. The cache fetches its card. The matcher finds a skill tagged warranty with text input and output, which the orchestrator needs. The loop publishes snapshot 2 with digest B.

Sessions started before 10:05 still resolve to digest A and never see the warranty agent. The first session after 10:05 triggers one build of root B. At 10:20 the returns team changes its card: a new skill, the same endpoint. The conditional GET returns 200 with new bytes, the hash differs, and the card is re-parsed. Digest C is published, and an alert reports that a skill list changed. At 10:30 the warranty agent's card endpoint times out. The cache serves the stale card and backs off. The loop marks the agent degraded after a configured number of failures and leaves it out of the next snapshot. Its health logic, not the cache, makes that decision.

Failure modes

Discovery fails in ways that are easy to miss in testing and painful in production.

  • Build-time fetches. Constructing RemoteA2AAgent without a card makes startup depend on every peer. Always pass the cached card.
  • Refresh storms. All entries expire together after a deploy, and every replica refetches every card. Add jitter to expiry times and keep refreshes single-flight.
  • Name collisions. Two cards share the name Support Agent, or a name is not a valid identifier. Derive local names from your configuration and reject duplicates when building a snapshot.
  • Flapping membership. A peer alternates between healthy and failing, and snapshots churn. Require several consecutive results before adding or removing an agent.
  • Trusting the source. A label or registry entry points at an unexpected host. Admit only allowlisted hosts over TLS, and check card signatures when your peers publish them.
  • Unbounded trees. Every discovered agent becomes a sub-agent, the routing prompt grows, and delegation accuracy drops. Cap the number of sub-agents per tree, and split by domain when the cap bites.
  • Snapshot leaks. Roots for old digests are never evicted. Reference-count them by pinned sessions.

Trade-offs

Static configuration against discovery. Static lists are reviewable and predictable. Discovery adapts to change but needs its own monitoring. Use static entries for critical dependencies and discovery for optional specialists. Pinned against live sessions. Pinning gives stable behaviour per conversation at the cost of memory for old trees and a delay before new agents are used. Live rebinding at each turn reacts faster but invites transfers to vanished agents. Short against long TTLs. Short TTLs notice changes quickly and load the peers; long TTLs are cheap but stale. Revalidation with ETags lets you have short TTLs cheaply, if peers support them. Central registry against well-known probing. A registry offers search, access control and selective disclosure, but it is a dependency and a single point of failure. Probing needs no shared infrastructure, but you must already know the domains.

What to do next

Use this checklist to add dynamic discovery to an ADK Java service.

  1. List your discovery sources, and mark which agents are critical and must stay in static configuration.
  2. Create a host allowlist, and enforce it in the resolver, not in each source.
  3. Implement the card cache with conditional GET, single-flight refresh, stale-on-error and jittered expiry.
  4. Define a reviewed tag vocabulary, and write the matcher against tags and modes.
  5. Compute snapshot digests from local names and card hashes; publish only when the digest changes.
  6. Build RemoteA2AAgent instances from cached cards with local names and reviewed descriptions.
  7. Pin each session to a digest; reference-count and evict old roots.
  8. Alert on skill-list changes, endpoint changes and repeated card-fetch failures.
  9. Cap sub-agents per tree, and test routing accuracy whenever the cap or vocabulary changes.
  10. Load-test a mass expiry after deploy to confirm refreshes stay single-flight.
Key takeaway: Run discovery as a background loop: merge sources behind a host allowlist, cache agent cards with conditional GETs, single-flight refresh and stale-on-error, and match on reviewed skill tags and modes. Publish immutable snapshots identified by digest, build one ADK root per digest from cached cards, and pin each session to the digest it started with.