An agent that delegates work to other agents inherits their failures. A peer that slows down holds your connections, a peer that loops spends your budget, a peer that returns something malformed or hostile feeds it straight into your next model call. In a monolith, one bad function crashes a request; in a mesh of agents, one bad peer can quietly drain every caller that depends on it, and the callers of those callers.
Failure isolation is the discipline of drawing walls so that a fault stays where it started. This article covers the walls that matter between A2A agents: per-peer bulkheads, a delegation budget that bounds time, depth and fan-out, treating remote output as untrusted data, quarantining inputs that keep killing peers, and cells on the serving side. Failing fast once a peer is known to be bad is the job of the circuit breaker, and putting things right afterwards belongs to error recovery; isolation is what limits the damage in between.
What can go wrong between agents
Start with a taxonomy, because each kind of fault spreads along a different path and needs a different wall. The A2A 1.0 specification gives you the vocabulary for outcomes: a task ends in TASK_STATE_COMPLETED, TASK_STATE_FAILED, TASK_STATE_CANCELED or TASK_STATE_REJECTED, the last meaning the agent decided not to do the work. It also defines InvalidAgentResponseError for responses that do not conform to the method. What the protocol cannot tell you is how a fault in one peer becomes a fault in you.
| Fault | How it spreads | Wall |
|---|---|---|
| Slow peer | Your in-flight calls pile up, holding sockets, memory and worker slots for every other peer | Per-peer bulkhead and deadline |
| Failing peer | Retries multiply load on it and latency for you | Retry budget, circuit breaker |
| Runaway delegation | Agents delegate to agents that delegate back, or fan out widely | Depth and fan-out budget |
| Malformed or hostile output | Remote text enters your prompt or tool arguments | Validation, data-only boundary |
| Poison input | One request crashes every peer it touches, then is retried | Fingerprint quarantine |
| Noisy tenant | One customer's burst exhausts a shared agent | Tenant cells and quotas |
Notice that only the second row is about errors in the usual sense. The most damaging faults are a peer that is merely slow and a peer that succeeds with bad content: neither raises an exception at the point where you could catch it.
The architecture: walls at every boundary
The diagram shows an orchestrator that delegates to three remote agents. Each remote peer has its own bulkhead, a fixed pool of concurrent calls. Each inbound task carries a budget that its children inherit in reduced form. Every remote result passes through a validator that turns it into structured data or rejects it, and inputs that repeatedly fail are fingerprinted and held in quarantine.
The essential property is that no resource is shared across peers without a cap. A single global connection pool or worker pool is the classic mistake: when one peer slows, its calls occupy the whole pool and every other peer appears to fail too, even though nothing is wrong with them.
Bulkheads per peer
A bulkhead caps how many delegations can be in flight to one peer and how many may wait for a slot. When both are full, it rejects immediately rather than queueing without limit, so the caller can degrade instead of hanging.
import asyncio, time
class BulkheadFull(Exception):
pass
class PeerBulkhead:
"""Caps concurrent and queued delegations to ONE remote agent."""
def __init__(self, name, max_inflight, max_waiting, max_wait_s):
self.name = name
self.sem = asyncio.Semaphore(max_inflight)
self.max_waiting = max_waiting
self.max_wait_s = max_wait_s
self.waiting = 0
async def run(self, coro_fn, deadline):
if self.waiting >= self.max_waiting:
raise BulkheadFull(self.name) # reject now, do not queue forever
self.waiting += 1
try:
budget = min(self.max_wait_s, deadline - time.monotonic())
await asyncio.wait_for(self.sem.acquire(), timeout=max(budget, 0))
except asyncio.TimeoutError:
raise BulkheadFull(self.name)
finally:
self.waiting -= 1
try:
remaining = deadline - time.monotonic()
return await asyncio.wait_for(coro_fn(), timeout=max(remaining, 0))
finally:
self.sem.release()
BULKHEADS = {
"flights": PeerBulkhead("flights", max_inflight=8, max_waiting=16, max_wait_s=0.5),
"hotels": PeerBulkhead("hotels", max_inflight=8, max_waiting=16, max_wait_s=0.5),
"fx": PeerBulkhead("fx", max_inflight=4, max_waiting=8, max_wait_s=0.2),
}Size each bulkhead from Little's law: concurrent calls equal arrival rate times latency. If you send hotels 4 requests per second at a normal latency of 1.5 seconds, you need about 6 slots; 8 leaves headroom. When hotels slows to 30 seconds, it would need 120 slots, and the bulkhead deliberately refuses to provide them. The short wait for a slot matters too: a caller that waits half a second and then gets BulkheadFull can return a partial answer, while one that waits indefinitely just moves the pile-up up a level. The general pattern, with thread-pool and semaphore variants, is in the bulkhead pattern article.
A delegation budget: deadline, depth and fan-out
Agents can call agents that call agents. Without a budget, a deep or cyclic delegation chain spends far more time and money than the original request was worth, and a single fan-out of ten that recurses three levels becomes a thousand remote tasks. Give every task a budget and give every child strictly less.
from dataclasses import dataclass, replace
@dataclass(frozen=True)
class Budget:
deadline: float # absolute, monotonic on this host; sent as remaining seconds
depth: int # hops left in the delegation tree
fanout: int # child delegations this node may still start
def child_budget(b: Budget, now: float, reserve_s: float = 0.3) -> Budget:
if b.depth <= 0:
raise PermissionError("delegation depth exhausted")
if b.fanout <= 0:
raise PermissionError("fan-out exhausted")
# keep a reserve so this node can still merge results and answer its own caller
return Budget(deadline=b.deadline - reserve_s, depth=b.depth - 1, fanout=4)
def to_metadata(b: Budget, now: float) -> dict:
# Application convention, not part of the A2A spec: carried in message metadata.
return {"x-budget": {"remaining_s": round(b.deadline - now, 3),
"depth": b.depth, "fanout": b.fanout}}Three choices make this work. Send the remaining time, not an absolute timestamp, because clocks on different hosts disagree; the receiver converts it to its own monotonic deadline. Hold back a reserve at each hop so the parent still has time to merge results and reply after a child times out. And reject when the depth or fan-out is spent rather than silently continuing. The A2A specification has no deadline or depth field, so this travels in message metadata as your own convention; agents that ignore it are still bounded by your local timeout, and agents you operate should enforce it on arrival. When a child times out, issue CancelTask for it, so it stops spending resources on an answer nobody will read. Deadline propagation and cancellation are covered further in timeout handling.
Remote output is data, never instructions
The fault most specific to agents is content. A remote agent's artifact may be malformed because of a bug, oversized because of a loop, or adversarial because the peer was fed a prompt injection by someone upstream. If you paste its text into your own model's prompt next to your tools, a compromise of that peer becomes a compromise of you.
MAX_ARTIFACT_BYTES = 256 * 1024
class RemoteTaskFailed(Exception): pass
class RemoteOutputRejected(Exception): pass
# Field names follow the JSON shape of a Task with artifacts made of parts; check the
# exact part representation your SDK version exposes before relying on "data".
def accept_remote_result(task: dict, schema_check) -> dict:
"""Turn a remote task into plain data, or raise. Never into instructions."""
state = task.get("status", {}).get("state")
if state != "TASK_STATE_COMPLETED":
raise RemoteTaskFailed(state) # FAILED, REJECTED, CANCELED...
parts = [p for a in task.get("artifacts", []) for p in a.get("parts", [])]
size = sum(len(str(p)) for p in parts)
if size > MAX_ARTIFACT_BYTES:
raise RemoteOutputRejected("artifact too large")
data = [p["data"] for p in parts if "data" in p]
if not data or not schema_check(data[0]):
raise RemoteOutputRejected("artifact does not match the agreed schema")
return data[0] # structured data onlyThe validator enforces four rules. Only a completed task counts as a result; failed, rejected and canceled states are handled as failures, not parsed. Size is capped before anything else touches the content. Structured data parts are preferred and checked against a schema agreed with the peer, such as a list of hotel offers with price, currency and identifier. And what comes out is data: your orchestrator formats it into a fixed template, and any free text from the peer is marked as quoted content in the prompt and never sits in a position where it can choose a tool.
Poison tasks and quarantine
Some inputs fail everywhere: a malformed document that crashes a parser, a request that sends a model into a loop until its token limit. Retry logic turns such an input into an amplifier, because each retry or re-delegation to another peer crashes that one too. Fingerprint the normalized input per skill, count failures across peers in a time window, and hold the fingerprint once it crosses a threshold.
import hashlib, time
class Quarantine:
def __init__(self, threshold=3, window_s=600, hold_s=3600):
self.fail = {} # fingerprint -> list of failure timestamps
self.held = {} # fingerprint -> release time
self.threshold, self.window_s, self.hold_s = threshold, window_s, hold_s
@staticmethod
def fingerprint(skill: str, normalized_input: str) -> str:
return hashlib.sha256(f"{skill}\x00{normalized_input}".encode()).hexdigest()[:16]
def check(self, fp):
if self.held.get(fp, 0) > time.time():
raise PermissionError(f"input {fp} quarantined")
def record_failure(self, fp, peer):
now = time.time()
hits = [t for t in self.fail.get(fp, []) if now - t < self.window_s] + [now]
self.fail[fp] = hits
if len(hits) >= self.threshold:
self.held[fp] = now + self.hold_s # alert a human with fp and peers involvedOnly count failures that look like the input's fault: crashes, task failures with an internal error, or timeouts at full budget. Do not count rejections or bulkhead refusals, which say more about the peer's capacity than the input. Quarantined inputs need a human path, a dead-letter queue with the fingerprint and peers involved, or you will silently drop legitimate but unusual work.
Isolation on the serving side: cells and tenants
If you operate an agent that others call, isolation also runs inward. Partition the deployment into cells, each a full copy of the agent with its own workers, model quota and state store, and pin each tenant to one cell by a stable hash. A tenant that sends a flood, or a payload that triggers a crash, takes out one cell, not the service. Apply per-tenant concurrency caps inside each cell as well, and use TASK_STATE_REJECTED or a protocol error early, at admission, rather than accepting a task you will later fail. Keep task state per cell, so that GetTask and SubscribeToTask for a tenant are served by the cell that owns the task.
Worked example: a slow hotels agent
A travel orchestrator handles 4 trip requests per second. Each trip delegates to a flights agent, a hotels agent and an FX agent, with a 6-second budget, depth 2 and fan-out 4. At 10:02 the hotels agent's backing service degrades and its latency rises from 1.5 to 30 seconds.
Without isolation, the orchestrator uses one pool of 64 outbound calls. Hotels calls arrive at 4 per second and never finish within a minute, so after about 16 seconds they occupy the whole pool. Flights and FX calls now wait for a slot, trips time out at 6 seconds with nothing, and dashboards show every peer failing.
With isolation, the hotels bulkhead fills its 8 slots within two seconds. New hotel delegations wait half a second, get BulkheadFull and the orchestrator returns flights and FX with a message that hotel options are temporarily unavailable. Each in-flight hotel call is cut at the trip deadline and cancelled. Flights and FX latency is unchanged. The circuit breaker opens a few seconds later and stops even the half-second wait.
Failure modes
| Symptom | Cause | Fix |
|---|---|---|
| Every peer looks down at once | Shared pool or event loop saturated by one slow peer | Per-peer bulkheads; never block the loop on a peer |
| Costs spike with no traffic change | Delegation cycle or fan-out recursion | Depth and fan-out budget, enforced on arrival |
| Children keep running after parent gave up | Timeout without cancellation | Issue CancelTask on deadline |
| Orchestrator takes actions nobody asked for | Peer text reached the prompt with tool access | Data-only boundary, schema validation |
| Same request fails across all peers repeatedly | Poison input retried and re-delegated | Fingerprint quarantine with dead-letter review |
| One tenant's burst degrades everyone | No cell or per-tenant cap on serving side | Cells pinned by tenant hash, per-tenant concurrency |
Trade-offs
Every wall wastes something. Bulkheads strand capacity: eight idle hotel slots cannot help a busy flights peer, so you provision more in total than a shared pool would need. Budgets cut off some requests that would have succeeded with a little more time. Cells multiply fixed costs and make cross-tenant features awkward. The return is predictability: the worst case of any single fault is bounded, and you can say in advance what fails when a given peer does. Start with bulkheads and budgets, which are cheap and catch the most common incidents, and add cells when your tenant count or blast-radius requirements justify them.
What to do next
- List every remote agent you delegate to and give each its own bulkhead, sized from rate times normal latency with headroom, and a short wait for a slot.
- Add a budget of remaining time, depth and fan-out to every outbound task, enforce it on inbound tasks, and cancel children when the deadline passes.
- Agree a structured artifact schema with each peer, cap artifact size, and stop free text from peers reaching any prompt position that can select a tool.
- Fingerprint failing inputs per skill and quarantine those that fail across peers, with a dead-letter path for review.
- If you serve agents to others, pin tenants to cells and cap per-tenant concurrency at admission.
- Pair this with the circuit breaker and load shedding articles, then run a game day that slows one peer to 30 seconds and checks that only its features degrade.