Once a multi-agent system has more than a handful of remote agents, teams add a registry: a service where agents announce their endpoint, version and capabilities, and where orchestrators look them up. The registry quickly becomes the most depended-on component in the fleet, and a registry outage can stop every agent from finding every other agent even though all of them are healthy. This article is about keeping that from happening.

Some scope first. ADK for Java does not ship a networked registry service of its own. Google Cloud offers a managed Agent Registry, and at the time of writing its documented ADK integration targets the Python ADK; with a managed registry the server side below is the provider's job, but the client-side rules still apply. The in-process registry pattern, a validated map of agent specs behind the runner, is covered in agent registry design, and the client side of finding remote agents, card caching, digests and session pinning is covered in dynamic agent discovery. The A2A protocol defines how an agent describes itself, with an agent card that is commonly served at /.well-known/agent-card.json, and lists curated registries as one discovery approach, but at the time of writing it does not standardise a registry API. So the registry discussed here is a service you build or adopt. The questions are how to replicate it, how registrations expire safely, how clients survive its outages, and how it recovers without being knocked over again.

Registry off the data path: replicated store, read replicas, cached clientsAgent instancesregister + heartbeat leaseRegistry write APIvalidates, fences by epochleaseQuorum store3 or 5 nodes across zonescommitRead replicassnapshot + watch per zonerevisionsADK orchestratorresolver with disk cachewatch / pullLast-known-goodserved when registry is downRemote agent endpointA2A calls go directdata path never touches the registryExpiry guardmass lease expiry freezes deletions instead of emptying the registry
Writes go through a quorum store; reads come from per-zone replicas into client caches; agent-to-agent calls never pass through the registry.

Keep the registry off the data path

The most important decision is where the registry sits. If every agent call asks the registry for an endpoint, the registry is on the data path and its availability multiplies into every request. A call that touches the registry, then a remote agent, then a model provider, each at 99.9%, is available about 99.7% of the time, and the registry contributed a third of the failures without doing any useful work.

Put it on the control path instead. Clients resolve names into a local snapshot, refresh that snapshot in the background, and call remote agents directly from it. A registry outage then means the snapshot stops updating, while existing traffic is unaffected. This property is often called static stability: when a dependency fails, the system keeps doing what it was doing rather than doing nothing. Every later section is a way of keeping that property intact, and the broader availability arithmetic for ADK services is in high availability for ADK Java.

Keep the registry off the data path

The most important decision is where the registry sits. If every agent call asks the registry for an endpoint, the registry is on the data path and its availability multiplies into every request. A call that touches the registry, then a remote agent, then a model provider, each at 99.9%, is available about 99.7% of the time, and the registry contributed a third of the failures without doing any useful work.

Put it on the control path instead. Clients resolve names into a local snapshot, refresh that snapshot in the background, and call remote agents directly from it. A registry outage then means the snapshot stops updating, while existing traffic is unaffected. This property is often called static stability: when a dependency fails, the system keeps doing what it was doing rather than doing nothing. Every later section is a way of keeping that property intact, and the broader availability arithmetic for ADK services is in high availability for ADK Java.

Registrations are leases with fencing

A registration is a lease, not a fact. An agent instance holds its entry only while it keeps renewing, so a crashed instance eventually disappears without anyone deleting it. Each record carries the fields that the rest of the design needs.

import java.time.Instant;

public record Registration(
        String agentName,       // logical name orchestrators resolve, e.g. "billing-agent"
        String instanceId,      // one per process; stable across heartbeats
        String endpoint,        // A2A base URL
        String cardDigest,      // sha-256 of the agent card this instance serves
        long epoch,             // incremented on every re-registration of this instance
        Instant leaseExpiresAt,
        String zone) {

    public boolean live(Instant now) {
        return now.isBefore(leaseExpiresAt);
    }
}

The epoch is a fencing token. If an instance pauses for a long garbage collection, its lease expires and a replacement registers with a higher epoch. When the paused instance wakes and sends a heartbeat carrying the old epoch, the registry rejects it instead of resurrecting a stale endpoint. The card digest lets clients notice that an endpoint's capabilities changed without fetching every card on every refresh.

The write path: quorum, leases and heartbeats

Registrations need a store that does not lose acknowledged writes and does not let two replicas disagree about who holds a lease. That points at a consensus-replicated store, such as etcd or a database with synchronous replication and conditional writes, deployed as three or five nodes across failure zones. Three nodes tolerate one failure, five tolerate two. Writes need a majority, so the write path stops when the majority is lost, and that is acceptable because reads do not depend on it.

The lease length is a trade-off between detection speed and false expiry. A common shape is a lease of a few times the heartbeat interval, so one or two lost heartbeats do not evict a healthy instance. Heartbeats must be jittered, or a fleet started together renews together for ever.

import java.net.URI;
import java.net.http.*;
import java.time.Duration;
import java.util.concurrent.*;

public final class LeaseKeeper {
    private static final Duration HEARTBEAT = Duration.ofSeconds(10);   // lease is set server-side to 30 s
    private final HttpClient http = HttpClient.newBuilder()
            .connectTimeout(Duration.ofSeconds(2)).build();
    private final ScheduledExecutorService timer = Executors.newSingleThreadScheduledExecutor();
    private final URI renewUri;
    private volatile long epoch;

    public LeaseKeeper(URI renewUri, long epoch) { this.renewUri = renewUri; this.epoch = epoch; }

    public void start() { schedule(); }

    private void schedule() {
        long jitterMs = ThreadLocalRandom.current().nextLong(-2000, 2000);
        timer.schedule(this::renew, HEARTBEAT.toMillis() + jitterMs, TimeUnit.MILLISECONDS);
    }

    private void renew() {
        HttpRequest req = HttpRequest.newBuilder(renewUri)
                .timeout(Duration.ofSeconds(3))
                .header("If-Match", Long.toString(epoch))   // fencing: stale epochs get 412
                .PUT(HttpRequest.BodyPublishers.noBody()).build();
        http.sendAsync(req, HttpResponse.BodyHandlers.discarding())
            .whenComplete((resp, err) -> {
                if (resp != null && resp.statusCode() == 412) {
                    reRegister();              // superseded: register again and take a new epoch
                }
                schedule();                    // errors just wait for the next jittered attempt
            });
    }

    private void reRegister() { /* POST the full registration; store the returned epoch */ }
}

The keeper treats registry errors as routine. A failed renewal is retried on the next tick, and the lease length absorbs a short outage. Only an explicit rejection of the epoch triggers re-registration.

The read path: replicas, watches and cached clients

Reads vastly outnumber writes, so serve them from replicas in every zone. Each replica holds a full snapshot in memory, tagged with the store revision it reflects, and follows changes through a watch stream. Clients ask for changes since a revision, so a reconnect costs a delta rather than a full download. If a replica falls too far behind and the store has compacted the history it needs, it reloads the full snapshot, and it must not serve reads that claim to be current while it does so.

On the client, the resolver keeps the last-known-good snapshot in memory and on local disk. The rule that matters is that a failure to refresh never empties the cache. Stale data is served with its age attached, and only a staleness beyond a hard limit, measured in hours rather than seconds, turns into an error. A process that restarts during a registry outage loads the disk copy and keeps working. The client-side details, including pinning a session to one snapshot so a conversation does not switch agent versions mid-turn, are in the dynamic discovery article linked above.

Health is not the registry's job alone. A lease says the process is alive and renewing; it does not say the agent answers correctly. Callers should still wrap each remote agent in timeouts and a circuit breaker, and eject an endpoint locally when it fails, without waiting for its lease to run out.

Guarding against mass expiry

The most damaging registry failure is not an outage but a mass deletion. If a network partition separates the registry from most agents, every lease expires on schedule, the registry faithfully removes healthy instances, and clients that trust it stop routing to agents that are working perfectly. Netflix's Eureka registry popularised a defence it calls self-preservation, and load balancers such as Envoy have a similar panic threshold: when too many endpoints look unhealthy at once, assume the observer is the problem.

import java.time.Instant;
import java.util.List;

final class ExpiryGuard {
    private static final double MAX_EXPIRY_FRACTION = 0.15;   // tune to your fleet's normal churn

    /** Returns the registrations that may be deleted on this sweep. */
    List<Registration> expirable(List<Registration> all, Instant now) {
        List<Registration> expired = all.stream().filter(r -> !r.live(now)).toList();
        if (all.isEmpty() || (double) expired.size() / all.size() <= MAX_EXPIRY_FRACTION) {
            return expired;
        }
        alert("lease expiry storm: " + expired.size() + " of " + all.size() + "; deletions frozen");
        return List.of();   // keep stale entries; callers' circuit breakers handle truly dead ones
    }

    private void alert(String msg) { System.err.println(msg); }
}

Freezing deletions trades a few stale endpoints for not erasing the fleet. That trade is safe only because callers have local ejection; without circuit breakers on the client side, the guard would send traffic to dead instances indefinitely.

Recovery and multiple regions

When the registry comes back after an outage, every agent re-registers and every client reconnects its watch at the same moment. A registry sized for steady state can fall over again under that herd, and come back, and fall over, in a loop. Four measures break the loop. Clients reconnect with exponential backoff and full jitter. Watches resume from their last revision instead of downloading the world. The write API applies admission control and sheds load with a clear retry-after signal rather than timing out. And heartbeats from instances the registry already knows are prioritised over brand-new registrations, because renewing an existing lease is cheap and keeps the snapshot accurate.

Across regions, avoid a single global quorum on any hot path. Run one registry cell per region, have agents register in their own region, and replicate a read-only view of other regions asynchronously for cross-region failover. A region that loses its link to the others keeps resolving local agents normally and serves its last copy of remote ones.

Worked example: a zone partition

A fleet runs 120 agent instances across three zones, with a five-node store spread two, two and one. A network fault isolates zone C, which holds one store node, one read replica and 40 instances.

In zone C, the replica loses its watch and keeps serving its snapshot. Agents in C cannot renew, and orchestrators in C continue calling agents in C directly from cached snapshots. In zones A and B, the store keeps its majority of four nodes and writes continue. Thirty seconds later the 40 zone C leases expire on the same sweep. That is a third of the fleet, far above the 15% guard, so deletions freeze and an alert fires. Orchestrators in A and B, whose calls into zone C now fail, eject those endpoints locally through their circuit breakers and route to the 80 instances they can reach.

When the link heals, zone C agents send heartbeats with their existing epochs. The guard's frozen entries are still present, so the renewals succeed without any re-registration, and the replica in C resumes its watch from its last revision. Circuit breakers in A and B close as probe calls succeed. Without the guard, the registry would have deleted 40 healthy instances and then absorbed 40 simultaneous re-registrations plus a full-snapshot reload in every client.

Failure modes

  • Registry on the data path. Per-call lookups turn a registry blip into a fleet-wide outage.
  • Cache cleared on error. A client that drops its snapshot when a refresh fails converts a control-plane outage into a data-plane one.
  • No fencing. A process resuming from a long pause resurrects its old endpoint after a replacement took over.
  • Synchronised heartbeats. Unjittered timers create load spikes and expiry waves.
  • Mass expiry. A partition between registry and agents deletes healthy instances.
  • Recovery herd. Simultaneous reconnects and full downloads knock the registry over again.
  • Global quorum. Cross-region consensus on writes makes every region depend on the slowest link.

Trade-offs

ChoiceGainCost
Short leaseDead instances vanish fastMore heartbeat load; false expiry on brief blips
Long staleness limit on clientsSurvives long outagesMay route to retired endpoints; needs local ejection
Expiry guardPartition cannot erase the fleetStale entries linger during real mass failure
Five-node storeTolerates two failuresHigher write latency and cost than three
Regional cellsNo cross-region hot dependencyGlobal view is eventually consistent

What to do next

  1. Check whether any agent call path performs a registry lookup per request, and move it to a background-refreshed snapshot.
  2. Persist the client snapshot to disk and test a cold start with the registry unreachable.
  3. Add an epoch to registrations and reject renewals that carry a stale one.
  4. Jitter heartbeats and set the lease to several heartbeat intervals.
  5. Add an expiry guard sized to your normal churn, with an alert when it trips.
  6. Make sure every remote agent call has a timeout and a circuit breaker, so local ejection does not wait for lease expiry.
  7. Run a game day: partition one zone, then restore it, and measure re-registration load and recovery time.
Key takeaway: Keep the registry on the control path so its outages freeze updates rather than stopping traffic. Model registrations as jittered, epoch-fenced leases in a quorum store, serve reads from per-zone replicas, and let clients keep a disk-backed last-known-good snapshot. Freeze deletions when too many leases expire at once, rely on client circuit breakers for real failures, and design recovery so reconnects are jittered and incremental.