Backpressure is the set of signals by which a component that is running out of capacity tells its callers to send less, and the reactions that make the callers actually do so. Inside one Java process you get it from bounded queues and reactive streams. Across an agent-to-agent (A2A) hop you get much less: the protocol carries tasks, messages and states, not credit windows. An orchestrator that delegates to a remote specialist therefore learns that the specialist is overloaded only indirectly, through HTTP errors, refused tasks or slow answers, and unless you build the reaction, it keeps sending.

This page lists the overload signals that can cross an A2A boundary, shows where each one surfaces in an ADK Java caller, and builds the code on both sides: an admission filter on the specialist that refuses work early, and a wrapper on the orchestrator that adapts its concurrency, honours deadlines and caps retries. The protocol statements were checked against the current A2A specification in October 2026; adk-java's a2a module on the main branch pins a2a-java 0.3.2.Final, so check how your versions surface each signal before relying on it.

Why agent chains amplify overload

Agent chains amplify overload in two ways. The first is retries: if each of three layers makes up to three attempts at a failed call, one user request can become 3 x 3 x 3 = 27 attempts at the bottom layer, arriving precisely when that layer is already failing. The second is cost per request: an agent task is not a cheap lookup but several model calls and tool calls, each consuming tokens against a provider quota. A specialist that accepts work it cannot finish wastes that quota on tasks whose callers have already timed out, which makes the overload last longer.

Good backpressure across agents therefore has three parts: the callee refuses early and cheaply, before spending model tokens; the caller reduces its sending rate when it sees refusals or rising latency; and everyone in the chain shares a deadline and a retry budget, so no layer works on requests that cannot succeed.

Backpressure across an A2A hop: what the callee can say and where the caller hears itOrchestrator (ADK Java)LlmAgent + remote sub-agentBackpressuredRemoteAgentdeadline, AIMD limit, budgetRemoteA2AAgenta2a-java clientAdmission filterslots + bounded waitA2A serverAgentExecutor + RunnerSpecialist agentmodel and toolsrequest + deadline503 / 429 + Retry-AfterSignals the caller can observeHTTP error -> A2AClientError (onError)REJECTED / FAILED task -> ordinary Eventrising latency -> only your own timer sees itCaller reactionsshrink concurrency (multiplicative)honour Retry-After, spend retry budgetskip the hop when the deadline is too close
The orchestrator's wrapper owns the reaction; the specialist's admission filter owns the refusal. The only shared vocabulary is HTTP status, task state and latency.

The signals that can cross the boundary

The A2A specification gives you less than you might expect. It says agents should rate-limit their operations and return appropriate errors when limits are exceeded, that servers may include retry guidance such as an HTTP Retry-After header, and that an authenticated extended agent card may describe rate limits and quotas. It defines no capacity field, no load-report message and no flow-control credit. Everything below is built from those parts.

SignalSent byMeaningHow the ADK Java caller sees it
HTTP 503 + Retry-AfterAdmission filter or proxyWhole service is at capacityAn exception through the run's onError, wrapped as A2AClientError
HTTP 429 + Retry-AfterRate limiterThis caller or tenant is over its quotaSame channel as 503; distinguish only if the status is exposed
Task state REJECTEDAgent logicThe agent declined this taskAn ordinary event, not an exception
Task state FAILEDAgent executorThe task started and brokeAn event with errorMessage set
Rising latencyNobodyQueues are growingOnly a timer you keep yourself

Two of these need care. Whether A2AClientError lets you read the HTTP status code and the Retry-After value depends on the a2a-java client version; if it does not, observe them where you control the HTTP layer, such as an egress proxy, and feed the result back. And REJECTED is listed as a terminal state, but the specification does not define what it implies about retrying, so treat it as a decision by the remote agent, not a capacity signal, unless you own both sides and document otherwise.

The callee: refuse early and cheaply

On the specialist, shed load before the request reaches the AgentExecutor, because after that point the agent starts calling the model. A servlet filter in front of the A2A endpoint holds a fixed number of task slots, lets a new request wait briefly for one, and otherwise refuses with 503 and a jittered Retry-After. Requests that do not start new work, such as fetching the agent card, reading a task or cancelling one, pass straight through: refusing a cancel during overload is exactly backwards.

public final class AdmissionFilter implements Filter {
  private final Semaphore slots;
  private final long maxWaitMs;

  public AdmissionFilter(int maxConcurrentTasks, long maxWaitMs) {
    this.slots = new Semaphore(maxConcurrentTasks, true);
    this.maxWaitMs = maxWaitMs;
  }

  @Override
  public void doFilter(ServletRequest req, ServletResponse res, FilterChain chain)
      throws IOException, ServletException {
    HttpServletRequest http = (HttpServletRequest) req;
    HttpServletResponse out = (HttpServletResponse) res;
    if (!startsNewWork(http)) { chain.doFilter(req, res); return; }   // card, get, cancel pass

    long remaining = Deadlines.remainingMs(http.getHeader("X-Deadline-Epoch-Ms"));
    if (remaining < Deadlines.MIN_USEFUL_MS) { refuse(out, 0, "deadline"); return; }

    boolean acquired = false;
    try {
      acquired = slots.tryAcquire(Math.min(maxWaitMs, remaining / 4), TimeUnit.MILLISECONDS);
      if (!acquired) { refuse(out, 1 + ThreadLocalRandom.current().nextInt(3), "capacity"); return; }
      chain.doFilter(req, res);
    } catch (InterruptedException e) {
      Thread.currentThread().interrupt();
      refuse(out, 1, "interrupted");
    } finally {
      if (acquired) slots.release();   // if streams complete asynchronously, release on completion
    }
  }

  private static void refuse(HttpServletResponse out, int retryAfterSeconds, String reason)
      throws IOException {
    out.setStatus(503);
    if (retryAfterSeconds > 0) out.setHeader("Retry-After", Integer.toString(retryAfterSeconds));
    out.setHeader("X-Shed-Reason", reason);
    out.getWriter().write("{\"error\":\"overloaded\",\"reason\":\"" + reason + "\"}");
  }
}

Three practical notes. With the JSON-RPC transport every call is a POST to one URL, so startsNewWork has to read the method name from the body; do that in a gateway that already parses it, or wrap the request so the body can be read twice. With streaming, the servlet call may return before the stream finishes, so release the slot when the stream completes, not in finally. And size the slot count from what the agent's model quota sustains, not from thread counts: if the provider allows 600 requests per minute and a task makes about six model calls lasting roughly 2 seconds each, the agent can sustain about 100 tasks per minute, and with a task taking about 12 seconds that is about 20 concurrent tasks.

Use 503 for capacity and 429 for per-tenant quota, so callers can tell whether slowing down will help everyone or only them. The X-Deadline-Epoch-Ms header is a convention of this page, not part of A2A; any agreed header or message metadata field works if both sides read it.

The caller: adaptive limits and a retry budget

On the orchestrator, the reaction is a concurrency limit per remote agent that adapts the way TCP congestion control does: grow slowly while calls succeed within a latency target, shrink by a constant factor on an overload signal. That converges on roughly what the specialist can absorb without anyone configuring it, and it shares capacity fairly between several orchestrator replicas, since each backs off independently. The retry budget sits beside it and allows retries only up to a fixed fraction of first attempts, which caps amplification however many layers retry.

/** Additive-increase, multiplicative-decrease concurrency limit for one remote agent. */
public final class AimdLimit {
  private final double min, max, backoff;
  private double limit;
  private int inFlight;

  public AimdLimit(double initial, double min, double max, double backoff) {
    this.limit = initial; this.min = min; this.max = max; this.backoff = backoff;
  }

  public synchronized boolean tryAcquire() {
    if (inFlight >= (int) limit) return false;
    inFlight++;
    return true;
  }

  public synchronized void onSuccess(long latencyMs, long targetMs) {
    inFlight--;
    if (latencyMs <= targetMs) limit = Math.min(max, limit + 1.0 / limit);  // about +1 per window
    else limit = Math.max(min, limit * 0.95);                               // slow is a soft signal
  }

  public synchronized void onOverload() { inFlight--; limit = Math.max(min, limit * backoff); }
  public synchronized void onOtherFailure() { inFlight--; }
  public synchronized double current() { return limit; }
}

/** Retries may spend at most `ratio` of first attempts, per caller process. */
public final class RetryBudget {
  private final double ratio, cap;
  private double tokens;
  public RetryBudget(double ratio, double cap) { this.ratio = ratio; this.cap = cap; }
  public synchronized void onFirstAttempt() { tokens = Math.min(cap, tokens + ratio); }
  public synchronized boolean tryRetry() { if (tokens < 1) return false; tokens -= 1; return true; }
}

The wrapper puts both around the remote agent and converts every outcome into explicit text the orchestrating model can read, the same technique the error-propagation page uses for failures. Wording matters: say the specialist is busy and should not be called again this turn, or the model will often call it again immediately.

// Your code, not an ADK API. BaseAgent's constructor differs between ADK releases; adapt it.
final class BackpressuredRemoteAgent extends BaseAgent {
  private final BaseAgent remote;          // the RemoteA2AAgent for one specialist
  private final AimdLimit limit;
  private final RetryBudget budget;
  private final long targetLatencyMs;

  @Override
  protected Flowable<Event> runAsyncImpl(InvocationContext ctx) {
    if (Deadlines.remainingMs(ctx) < Deadlines.MIN_USEFUL_MS) {
      return Flowable.just(note(ctx, "REMOTE_SKIPPED: not enough time left in this turn"));
    }
    budget.onFirstAttempt();
    if (!limit.tryAcquire()) {
      return Flowable.just(note(ctx, "REMOTE_BUSY: local limit reached; answer without it"));
    }
    long start = System.nanoTime();
    AtomicReference<Outcome> outcome = new AtomicReference<>(Outcome.OK);
    return remote.runAsync(ctx)
        .doOnNext(ev -> ev.errorMessage().ifPresent(m -> outcome.set(Outcome.FAILED)))
        .onErrorResumeNext(t -> {
          boolean overload = Overload.isOverload(t);   // 429/503 if your client exposes status
          outcome.set(overload ? Outcome.OVERLOAD : Outcome.FAILED);
          return Flowable.just(note(ctx, overload
              ? "REMOTE_BUSY: specialist is shedding load; do not call it again this turn"
              : "REMOTE_UNAVAILABLE: no task was confirmed"));
        })
        .doFinally(() -> {
          long ms = (System.nanoTime() - start) / 1_000_000;
          switch (outcome.get()) {
            case OK -> limit.onSuccess(ms, targetLatencyMs);
            case OVERLOAD -> limit.onOverload();
            case FAILED -> limit.onOtherFailure();
          }
        });
  }
}

Deadlines across hops

Deadlines are what stop doomed work. The orchestrator computes an absolute deadline when the user's request arrives, for example arrival plus 30 seconds, and passes it to every hop. Each hop subtracts nothing and invents nothing; it simply refuses to start when the remaining time is below the minimum useful amount for its kind of task, and caps its own internal waits to a fraction of the remainder, as the admission filter does with its queue wait. Use absolute epoch time rather than a relative timeout so that time spent in queues and network is counted once, and keep server clocks synchronised. Whether RemoteA2AAgent lets you add a header is unconfirmed; if not, inject it at an egress proxy.

Where a retry is allowed, it must pass three checks: the budget has a token, the deadline leaves time for another full attempt, and Retry-After, if known, falls before the deadline. If any check fails, return the busy note and let the orchestrating model answer with what it has.

Worked example: a spike on a research specialist

An orchestrator with four replicas delegates research to a specialist sized at 20 slots. Each replica's limit starts at 10, so together they can send 40 concurrent tasks. A traffic spike arrives. The specialist fills its 20 slots, holds a few requests for the bounded wait and refuses the rest with 503. Each replica that sees a refusal multiplies its limit by 0.7, from 10 to 7, then to about 5. Within a few seconds the replicas together send roughly 20, the refusals stop, latencies return under target, and the limits creep upward again by about one per window until the next refusal. The specialist's model quota was never exceeded and no task it accepted was abandoned.

Without the wrapper, the same spike looks different: each replica keeps 10 tasks in flight, the overflow is refused, and each layer's retries add more load. Users waiting on refused requests see errors after the full timeout instead of a fast answer that the research step was skipped.

Observability

  • Per remote agent: current AIMD limit, in-flight count, overload refusals and retry-budget exhaustion. A limit pinned at its minimum means the specialist is undersized.
  • On the specialist: slot utilisation, queue wait, refusals by reason, and tasks started with less than the minimum useful time left (should be zero).
  • Across the chain: the share of user turns in which any hop was skipped, which is the user-visible cost of shedding.

Failure modes

  • Refusing after spending. Admission checks placed inside the agent, after the first model call, shed load too late.
  • Leaked slots. Streaming responses release the semaphore before the stream ends, so concurrency silently exceeds the limit.
  • Synchronised retries. A fixed Retry-After makes every caller return at the same instant; add jitter.
  • Model-driven retries. The orchestrating model calls the busy specialist again because the note did not tell it not to.
  • Overload hidden as failure. If 503 and an agent bug look the same to the caller, the limit either never shrinks or shrinks on bugs. Classify them separately.

Trade-offs

ChoiceBenefitCost
Adaptive limit versus fixed limitTracks real capacity, no tuning per deploymentOscillates; needs a sensible latency target
Short queue wait at the calleeAbsorbs brief burstsAdds latency to refused requests
Skip the hop versus fail the turnUser still gets an answerThe answer is less complete; tell the user
Deadline header conventionStops doomed work everywhereBoth sides must adopt it

What to do next

  1. For each remote agent, find out how your a2a-java version surfaces HTTP status and Retry-After, and write the result down.
  2. Add an admission filter to each specialist sized from its model quota, refusing with 503 and jitter; let card, get and cancel calls through.
  3. Wrap each remote agent with an adaptive limit, a retry budget and busy notes the model can act on.
  4. Introduce an absolute deadline at the edge and refuse work below the minimum useful time at every hop.
  5. Load-test with a deliberately undersized specialist and check that refusals stop and no accepted task is abandoned.
  6. Continue with A2A error propagation, the ADK Java circuit breaker, tool retry and backoff, rate limiting parallel agents and A2A load shedding.
Key takeaway: A2A carries no flow-control messages, so backpressure between agents is built from HTTP status codes with Retry-After, task states and latency. Refuse new tasks on the specialist before it spends model tokens, with a slot-based admission filter that lets cancel and read calls through. On the orchestrator, wrap each remote agent with an additive-increase, multiplicative-decrease concurrency limit, a retry budget that caps amplification, and busy notes the model can act on. Propagate an absolute deadline so no hop starts work that cannot finish, and check how your a2a-java version exposes status codes before you depend on them.