An agent that delegates work to other agents inherits their failures. A peer that slows down holds your connections, a peer that loops spends your budget, a peer that returns something malformed or hostile feeds it straight into your next model call. In a monolith, one bad function crashes a request; in a mesh of agents, one bad peer can quietly drain every caller that depends on it, and the callers of those callers.

Failure isolation is the discipline of drawing walls so that a fault stays where it started. This article covers the walls that matter between A2A agents: per-peer bulkheads, a delegation budget that bounds time, depth and fan-out, treating remote output as untrusted data, quarantining inputs that keep killing peers, and cells on the serving side. Failing fast once a peer is known to be bad is the job of the circuit breaker, and putting things right afterwards belongs to error recovery; isolation is what limits the damage in between.

Advertisement

What can go wrong between agents

Start with a taxonomy, because each kind of fault spreads along a different path and needs a different wall. The A2A 1.0 specification gives you the vocabulary for outcomes: a task ends in TASK_STATE_COMPLETED, TASK_STATE_FAILED, TASK_STATE_CANCELED or TASK_STATE_REJECTED, the last meaning the agent decided not to do the work. It also defines InvalidAgentResponseError for responses that do not conform to the method. What the protocol cannot tell you is how a fault in one peer becomes a fault in you.

FaultHow it spreadsWall
Slow peerYour in-flight calls pile up, holding sockets, memory and worker slots for every other peerPer-peer bulkhead and deadline
Failing peerRetries multiply load on it and latency for youRetry budget, circuit breaker
Runaway delegationAgents delegate to agents that delegate back, or fan out widelyDepth and fan-out budget
Malformed or hostile outputRemote text enters your prompt or tool argumentsValidation, data-only boundary
Poison inputOne request crashes every peer it touches, then is retriedFingerprint quarantine
Noisy tenantOne customer's burst exhausts a shared agentTenant cells and quotas

Notice that only the second row is about errors in the usual sense. The most damaging faults are a peer that is merely slow and a peer that succeeds with bad content: neither raises an exception at the point where you could catch it.

The architecture: walls at every boundary

The diagram shows an orchestrator that delegates to three remote agents. Each remote peer has its own bulkhead, a fixed pool of concurrent calls. Each inbound task carries a budget that its children inherit in reduced form. Every remote result passes through a validator that turns it into structured data or rejects it, and inputs that repeatedly fail are fingerprinted and held in quarantine.

Inbound tasktenant cell AOrchestratorbudget: deadline,depth, fan-outBulkheadflights: 8 slotsBulkheadhotels: 8/8 fullBulkheadfx: 4 slotsFlights agentremoteHotels agentslow: 30 sFX agentremoteOutput validatorschema, size, no toolsvalid data onlyQuarantinepoison fingerprintsOne slow peer fills only its own bulkhead; flights and FX keep their slots
Per-peer bulkheads, an inherited delegation budget, an output validator and a quarantine. The hotels agent has slowed to 30 seconds and filled its own eight slots; flights and FX are untouched.

The essential property is that no resource is shared across peers without a cap. A single global connection pool or worker pool is the classic mistake: when one peer slows, its calls occupy the whole pool and every other peer appears to fail too, even though nothing is wrong with them.

Advertisement

Bulkheads per peer

A bulkhead caps how many delegations can be in flight to one peer and how many may wait for a slot. When both are full, it rejects immediately rather than queueing without limit, so the caller can degrade instead of hanging.

import asyncio, time

class BulkheadFull(Exception):
    pass

class PeerBulkhead:
    """Caps concurrent and queued delegations to ONE remote agent."""
    def __init__(self, name, max_inflight, max_waiting, max_wait_s):
        self.name = name
        self.sem = asyncio.Semaphore(max_inflight)
        self.max_waiting = max_waiting
        self.max_wait_s = max_wait_s
        self.waiting = 0

    async def run(self, coro_fn, deadline):
        if self.waiting >= self.max_waiting:
            raise BulkheadFull(self.name)            # reject now, do not queue forever
        self.waiting += 1
        try:
            budget = min(self.max_wait_s, deadline - time.monotonic())
            await asyncio.wait_for(self.sem.acquire(), timeout=max(budget, 0))
        except asyncio.TimeoutError:
            raise BulkheadFull(self.name)
        finally:
            self.waiting -= 1
        try:
            remaining = deadline - time.monotonic()
            return await asyncio.wait_for(coro_fn(), timeout=max(remaining, 0))
        finally:
            self.sem.release()

BULKHEADS = {
    "flights": PeerBulkhead("flights", max_inflight=8, max_waiting=16, max_wait_s=0.5),
    "hotels":  PeerBulkhead("hotels",  max_inflight=8, max_waiting=16, max_wait_s=0.5),
    "fx":      PeerBulkhead("fx",      max_inflight=4, max_waiting=8,  max_wait_s=0.2),
}

Size each bulkhead from Little's law: concurrent calls equal arrival rate times latency. If you send hotels 4 requests per second at a normal latency of 1.5 seconds, you need about 6 slots; 8 leaves headroom. When hotels slows to 30 seconds, it would need 120 slots, and the bulkhead deliberately refuses to provide them. The short wait for a slot matters too: a caller that waits half a second and then gets BulkheadFull can return a partial answer, while one that waits indefinitely just moves the pile-up up a level. The general pattern, with thread-pool and semaphore variants, is in the bulkhead pattern article.

A delegation budget: deadline, depth and fan-out

Agents can call agents that call agents. Without a budget, a deep or cyclic delegation chain spends far more time and money than the original request was worth, and a single fan-out of ten that recurses three levels becomes a thousand remote tasks. Give every task a budget and give every child strictly less.

from dataclasses import dataclass, replace

@dataclass(frozen=True)
class Budget:
    deadline: float      # absolute, monotonic on this host; sent as remaining seconds
    depth: int           # hops left in the delegation tree
    fanout: int          # child delegations this node may still start

def child_budget(b: Budget, now: float, reserve_s: float = 0.3) -> Budget:
    if b.depth <= 0:
        raise PermissionError("delegation depth exhausted")
    if b.fanout <= 0:
        raise PermissionError("fan-out exhausted")
    # keep a reserve so this node can still merge results and answer its own caller
    return Budget(deadline=b.deadline - reserve_s, depth=b.depth - 1, fanout=4)

def to_metadata(b: Budget, now: float) -> dict:
    # Application convention, not part of the A2A spec: carried in message metadata.
    return {"x-budget": {"remaining_s": round(b.deadline - now, 3),
                         "depth": b.depth, "fanout": b.fanout}}

Three choices make this work. Send the remaining time, not an absolute timestamp, because clocks on different hosts disagree; the receiver converts it to its own monotonic deadline. Hold back a reserve at each hop so the parent still has time to merge results and reply after a child times out. And reject when the depth or fan-out is spent rather than silently continuing. The A2A specification has no deadline or depth field, so this travels in message metadata as your own convention; agents that ignore it are still bounded by your local timeout, and agents you operate should enforce it on arrival. When a child times out, issue CancelTask for it, so it stops spending resources on an answer nobody will read. Deadline propagation and cancellation are covered further in timeout handling.

Remote output is data, never instructions

The fault most specific to agents is content. A remote agent's artifact may be malformed because of a bug, oversized because of a loop, or adversarial because the peer was fed a prompt injection by someone upstream. If you paste its text into your own model's prompt next to your tools, a compromise of that peer becomes a compromise of you.

MAX_ARTIFACT_BYTES = 256 * 1024

class RemoteTaskFailed(Exception): pass
class RemoteOutputRejected(Exception): pass

# Field names follow the JSON shape of a Task with artifacts made of parts; check the
# exact part representation your SDK version exposes before relying on "data".
def accept_remote_result(task: dict, schema_check) -> dict:
    """Turn a remote task into plain data, or raise. Never into instructions."""
    state = task.get("status", {}).get("state")
    if state != "TASK_STATE_COMPLETED":
        raise RemoteTaskFailed(state)              # FAILED, REJECTED, CANCELED...
    parts = [p for a in task.get("artifacts", []) for p in a.get("parts", [])]
    size = sum(len(str(p)) for p in parts)
    if size > MAX_ARTIFACT_BYTES:
        raise RemoteOutputRejected("artifact too large")
    data = [p["data"] for p in parts if "data" in p]
    if not data or not schema_check(data[0]):
        raise RemoteOutputRejected("artifact does not match the agreed schema")
    return data[0]                                 # structured data only

The validator enforces four rules. Only a completed task counts as a result; failed, rejected and canceled states are handled as failures, not parsed. Size is capped before anything else touches the content. Structured data parts are preferred and checked against a schema agreed with the peer, such as a list of hotel offers with price, currency and identifier. And what comes out is data: your orchestrator formats it into a fixed template, and any free text from the peer is marked as quoted content in the prompt and never sits in a position where it can choose a tool.

Poison tasks and quarantine

Some inputs fail everywhere: a malformed document that crashes a parser, a request that sends a model into a loop until its token limit. Retry logic turns such an input into an amplifier, because each retry or re-delegation to another peer crashes that one too. Fingerprint the normalized input per skill, count failures across peers in a time window, and hold the fingerprint once it crosses a threshold.

import hashlib, time

class Quarantine:
    def __init__(self, threshold=3, window_s=600, hold_s=3600):
        self.fail = {}          # fingerprint -> list of failure timestamps
        self.held = {}          # fingerprint -> release time
        self.threshold, self.window_s, self.hold_s = threshold, window_s, hold_s

    @staticmethod
    def fingerprint(skill: str, normalized_input: str) -> str:
        return hashlib.sha256(f"{skill}\x00{normalized_input}".encode()).hexdigest()[:16]

    def check(self, fp):
        if self.held.get(fp, 0) > time.time():
            raise PermissionError(f"input {fp} quarantined")

    def record_failure(self, fp, peer):
        now = time.time()
        hits = [t for t in self.fail.get(fp, []) if now - t < self.window_s] + [now]
        self.fail[fp] = hits
        if len(hits) >= self.threshold:
            self.held[fp] = now + self.hold_s   # alert a human with fp and peers involved

Only count failures that look like the input's fault: crashes, task failures with an internal error, or timeouts at full budget. Do not count rejections or bulkhead refusals, which say more about the peer's capacity than the input. Quarantined inputs need a human path, a dead-letter queue with the fingerprint and peers involved, or you will silently drop legitimate but unusual work.

Isolation on the serving side: cells and tenants

If you operate an agent that others call, isolation also runs inward. Partition the deployment into cells, each a full copy of the agent with its own workers, model quota and state store, and pin each tenant to one cell by a stable hash. A tenant that sends a flood, or a payload that triggers a crash, takes out one cell, not the service. Apply per-tenant concurrency caps inside each cell as well, and use TASK_STATE_REJECTED or a protocol error early, at admission, rather than accepting a task you will later fail. Keep task state per cell, so that GetTask and SubscribeToTask for a tenant are served by the cell that owns the task.

Worked example: a slow hotels agent

A travel orchestrator handles 4 trip requests per second. Each trip delegates to a flights agent, a hotels agent and an FX agent, with a 6-second budget, depth 2 and fan-out 4. At 10:02 the hotels agent's backing service degrades and its latency rises from 1.5 to 30 seconds.

Without isolation, the orchestrator uses one pool of 64 outbound calls. Hotels calls arrive at 4 per second and never finish within a minute, so after about 16 seconds they occupy the whole pool. Flights and FX calls now wait for a slot, trips time out at 6 seconds with nothing, and dashboards show every peer failing.

With isolation, the hotels bulkhead fills its 8 slots within two seconds. New hotel delegations wait half a second, get BulkheadFull and the orchestrator returns flights and FX with a message that hotel options are temporarily unavailable. Each in-flight hotel call is cut at the trip deadline and cancelled. Flights and FX latency is unchanged. The circuit breaker opens a few seconds later and stops even the half-second wait.

Failure modes

SymptomCauseFix
Every peer looks down at onceShared pool or event loop saturated by one slow peerPer-peer bulkheads; never block the loop on a peer
Costs spike with no traffic changeDelegation cycle or fan-out recursionDepth and fan-out budget, enforced on arrival
Children keep running after parent gave upTimeout without cancellationIssue CancelTask on deadline
Orchestrator takes actions nobody asked forPeer text reached the prompt with tool accessData-only boundary, schema validation
Same request fails across all peers repeatedlyPoison input retried and re-delegatedFingerprint quarantine with dead-letter review
One tenant's burst degrades everyoneNo cell or per-tenant cap on serving sideCells pinned by tenant hash, per-tenant concurrency

Trade-offs

Every wall wastes something. Bulkheads strand capacity: eight idle hotel slots cannot help a busy flights peer, so you provision more in total than a shared pool would need. Budgets cut off some requests that would have succeeded with a little more time. Cells multiply fixed costs and make cross-tenant features awkward. The return is predictability: the worst case of any single fault is bounded, and you can say in advance what fails when a given peer does. Start with bulkheads and budgets, which are cheap and catch the most common incidents, and add cells when your tenant count or blast-radius requirements justify them.

What to do next

  1. List every remote agent you delegate to and give each its own bulkhead, sized from rate times normal latency with headroom, and a short wait for a slot.
  2. Add a budget of remaining time, depth and fan-out to every outbound task, enforce it on inbound tasks, and cancel children when the deadline passes.
  3. Agree a structured artifact schema with each peer, cap artifact size, and stop free text from peers reaching any prompt position that can select a tool.
  4. Fingerprint failing inputs per skill and quarantine those that fail across peers, with a dead-letter path for review.
  5. If you serve agents to others, pin tenants to cells and cap per-tenant concurrency at admission.
  6. Pair this with the circuit breaker and load shedding articles, then run a game day that slows one peer to 30 seconds and checks that only its features degrade.
Key takeaway: In an agent-to-agent system the worst faults are slow peers and peers that succeed with bad content, and both spread through shared resources and trusted text. Give each remote agent its own bounded bulkhead, pass a shrinking budget of time, depth and fan-out down every delegation and cancel children when it runs out, treat remote artifacts as validated data and never as instructions, quarantine inputs that fail everywhere, and pin tenants to cells on the serving side.