When one agent calls another too often, the callee has to say no. How it says no decides what happens next. A clear rejection tells the caller whose limit was hit, how long to wait and whether the request did any work. A vague one, such as a bare 500, makes the caller guess, and guessing callers either give up on work that would have succeeded a few seconds later or retry immediately and make the overload worse.

The A2A specification asks agents to rate limit but does not define a rate-limit error. This page sets out exactly what A2A 1.0 says, proposes one rejection shape that works across its three bindings, and gives server and client code for it. How to build the limiter itself, token buckets and distributed counters, is covered in A2A rate limiting. The full list of A2A error codes is in A2A Error Codes, in depth.

Advertisement

What A2A 1.0 says, and what it leaves to you

The specification touches rate limiting in two places. Its security section says agents SHOULD implement rate limiting on all operations, SHOULD return appropriate error responses when limits are exceeded, and MAY use different limits per operation or tier. Its general error-handling section lists "rate limit exceeded" as an example scenario under System Errors, next to database failures and downstream timeouts, with example codes HTTP 500 or 503, gRPC INTERNAL or UNAVAILABLE, and JSON-RPC -32603. The same section says servers MAY include retry guidance, for example a Retry-After header.

Three things follow. There is no A2A-specific rate-limit error, so do not invent one: codes -32010 and above are unassigned and a later revision may take them. Every error response must carry a code, a message and optional typed details using the google.rpc error model, so retry guidance belongs in those details. And because the codes in the spec are examples, you have room to choose more precise ones. This page recommends HTTP 429 and gRPC RESOURCE_EXHAUSTED for a caller's quota. That is a recommendation, not a spec rule, and clients should still accept the spec's example codes from agents that follow them literally.

What a rejection must tell the caller

Before choosing codes, decide what the caller needs to know to act correctly.

  • Is it me or everyone? A caller that exceeded its own quota should slow down; the agent is healthy for everyone else. An agent that is overloaded is failing for all callers, and the caller should consider another agent or a degraded path. These are different situations and should have different codes: 429 or RESOURCE_EXHAUSTED for quota, 503 or UNAVAILABLE for overload.
  • When can I try again? The server knows when the window resets or the bucket refills; the client does not. Send that delay explicitly rather than making every client guess with exponential backoff.
  • Which limit? Agents often have several limits: per caller, per skill, per tenant, concurrent tasks. Name the one that was hit so the caller can tell a burst from a daily cap.
  • Did anything happen? A rate-limit rejection must be decided at admission, before a task is created or any side effect starts. Then the answer is always no, and the retry is safe. Say so in your documentation and make it true in code.
Admission path: decide, and shape the rejection, before any A2A work startsRequestany bindingAuthenticatecaller identityCaller quota?token bucket per callerAgent overloaded?concurrency, queue depthokQuota exceeded429 / RESOURCE_EXHAUSTEDOverload503 / UNAVAILABLEoveryesAdmittask created, stream opensnoEvery rejection carriesRetryInfo delay and Retry-AfterErrorInfo reason in your own domainQuotaFailure naming the exhausted quotaAfter a stream has opened with 200, a limit can no longer become a 429; it has to be handled inside the task.
Figure 1. Limits are enforced at admission, after authentication so the caller is known, and before any task exists. Quota and overload produce different codes; both carry machine-readable retry guidance.
Advertisement

The format on the HTTP+JSON binding

The HTTP+JSON binding returns errors as google.rpc.Status JSON under an error key, with the HTTP status mirrored in error.code, and the media type application/a2a+json. A rate-limit rejection then looks like the response below. It uses three well-known detail types. ErrorInfo carries a stable reason under your own domain: the a2a-protocol.org domain is reserved for A2A's own errors. RetryInfo carries the delay as a protobuf Duration string. QuotaFailure names the subject and the quota that was exceeded.

HTTP/1.1 429 Too Many Requests
Content-Type: application/a2a+json
Retry-After: 12
RateLimit-Policy: "per-caller";q=60;w=60
RateLimit: "per-caller";r=0;t=12

{
  "error": {
    "code": 429,
    "status": "RESOURCE_EXHAUSTED",
    "message": "Caller orchestrator-7 exceeded 60 SendMessage calls per minute",
    "details": [
      {"@type": "type.googleapis.com/google.rpc.ErrorInfo",
       "reason": "RATE_LIMIT_EXCEEDED", "domain": "agents.example.com",
       "metadata": {"policy": "per-caller", "scope": "caller", "method": "SendMessage"}},
      {"@type": "type.googleapis.com/google.rpc.RetryInfo", "retryDelay": "12s"},
      {"@type": "type.googleapis.com/google.rpc.QuotaFailure",
       "violations": [{"subject": "caller:orchestrator-7",
                       "description": "60 requests per 60 s"}]}
    ]
  }
}

The headers repeat the delay for clients and proxies that never read bodies. Retry-After is the standard field and may be either a number of seconds or an HTTP date; send seconds, because they do not depend on clocks agreeing. RateLimit-Policy and RateLimit come from the IETF HTTPAPI working group's draft, revision 11 at the time of writing, which is still an Internet-Draft. In it, q is the quota, w the window in seconds, r the remaining quota and t the seconds until it resets. The draft may still change, and many clients only know older X-RateLimit-* fields, so treat these as hints and keep the body authoritative. Sending RateLimit on successful responses too lets callers slow down before they are rejected.

gRPC and JSON-RPC

On gRPC, return status RESOURCE_EXHAUSTED for quota and UNAVAILABLE for overload, with the same three details in google.rpc.Status.details. One caution: gRPC libraries also produce RESOURCE_EXHAUSTED locally, for example when a message exceeds the size limit. A client must not treat every RESOURCE_EXHAUSTED as a rate limit; it should check for your ErrorInfo reason, and treat the status without it as a non-retryable error.

The JSON-RPC binding has no slot for an HTTP-style status inside the error object, and A2A 1.0 does not say which HTTP status a JSON-RPC error response should travel with. Use -32603, which is the spec's own example, and put the meaning in error.data, which A2A defines as an array of typed objects rather than JSON-RPC's free-form value. Also sending HTTP 429 helps gateways and generic HTTP clients, but some JSON-RPC libraries treat any non-200 as a transport failure and never parse the body. Send 429 to callers you control; otherwise send 200 with the body and keep Retry-After as a header. Either way the body must carry everything.

HTTP/1.1 429 Too Many Requests
Content-Type: application/json
Retry-After: 12

{
  "jsonrpc": "2.0",
  "id": "req-41",
  "error": {
    "code": -32603,
    "message": "Rate limit exceeded",
    "data": [
      {"@type": "type.googleapis.com/google.rpc.ErrorInfo",
       "reason": "RATE_LIMIT_EXCEEDED", "domain": "agents.example.com",
       "metadata": {"policy": "per-caller", "scope": "caller"}},
      {"@type": "type.googleapis.com/google.rpc.RetryInfo", "retryDelay": "12s"}
    ]
  }
}

Server code: one decision, three encoders

Make the limiter return a single rejection value and let each binding encode it, then test that every encoding carries the same reason and delay.

from dataclasses import dataclass

DOMAIN = "agents.example.com"          # yours, never "a2a-protocol.org"

@dataclass
class Rejection:
    kind: str            # "quota" or "overload"
    retry_after_s: int
    policy: str
    subject: str
    message: str

def details(r: Rejection):
    out = [{"@type": "type.googleapis.com/google.rpc.ErrorInfo",
            "reason": "RATE_LIMIT_EXCEEDED" if r.kind == "quota" else "AGENT_OVERLOADED",
            "domain": DOMAIN,
            "metadata": {"policy": r.policy, "scope": "caller" if r.kind == "quota" else "agent"}},
           {"@type": "type.googleapis.com/google.rpc.RetryInfo",
            "retryDelay": f"{r.retry_after_s}s"}]
    if r.kind == "quota":
        out.append({"@type": "type.googleapis.com/google.rpc.QuotaFailure",
                    "violations": [{"subject": r.subject, "description": r.policy}]})
    return out

def http_status(r):  return 429 if r.kind == "quota" else 503
def grpc_status(r):  return "RESOURCE_EXHAUSTED" if r.kind == "quota" else "UNAVAILABLE"

def encode_rest(r):
    body = {"error": {"code": http_status(r), "status": grpc_status(r),
                      "message": r.message, "details": details(r)}}
    return http_status(r), {"Retry-After": str(r.retry_after_s)}, body

def encode_jsonrpc(r, req_id):
    body = {"jsonrpc": "2.0", "id": req_id,
            "error": {"code": -32603, "message": r.message, "data": details(r)}}
    return http_status(r), {"Retry-After": str(r.retry_after_s)}, body

Round the delay up to whole seconds, and never send zero: a zero delay invites an immediate retry that will be rejected again. Compute it from the limiter's actual state, the time until the bucket holds one token, rather than a constant, so callers that obey it succeed on their first retry.

Client code: wait the right amount

The client's job is to turn any of these shapes into a delay. Prefer the most precise source: RetryInfo from the body, then Retry-After, then the draft RateLimit header. Only fall back to exponential backoff when the server gave no hint. Add jitter in both cases: if a hundred callers were rejected in the same second and all wait exactly twelve seconds, they all return in the same second.

import random, re
from datetime import datetime, timezone
from email.utils import parsedate_to_datetime

def retry_after_seconds(headers, details, now=None):
    """Prefer RetryInfo, then Retry-After (seconds or HTTP-date), then RateLimit t=."""
    for d in details or []:
        if d.get("@type", "").endswith("google.rpc.RetryInfo"):
            m = re.fullmatch(r"(\d+(?:\.\d+)?)s", d.get("retryDelay", ""))
            if m:
                return float(m.group(1))
    ra = headers.get("Retry-After")
    if ra:
        if ra.strip().isdigit():
            return float(ra)
        try:
            when = parsedate_to_datetime(ra)
            now = now or datetime.now(timezone.utc)
            return max(0.0, (when - now).total_seconds())
        except (TypeError, ValueError):
            pass
    m = re.search(r";\s*t=(\d+)", headers.get("RateLimit", ""))
    return float(m.group(1)) if m else None

def next_delay(attempt, hinted, base=0.5, cap=60.0):
    if hinted is not None:                       # the server knows its window
        return min(cap, hinted) + random.uniform(0, 1.0)
    return random.uniform(0, min(cap, base * 2 ** attempt))   # full jitter

Three rules make retries safe. Retry with the same messageId, so that an agent that did admit an earlier copy can deduplicate it. Cap the total time spent retrying with a deadline taken from the caller's own budget, and fail the parent task cleanly when it runs out. And share the delay across concurrent requests to the same agent: if one call was told to wait twelve seconds, the next ten should not each find that out separately. A2A Error Recovery covers what to do when the retries run out.

Streaming and push notifications

A streaming call returns HTTP 200 with text/event-stream and then sends a Task followed by update events until a terminal state. Once that 200 is sent, there is no way to turn the response into a 429. So check every limit that can reject the call, including concurrent-stream limits, before writing the first byte. If the agent later hits a limit of its own, such as a model provider's quota, that is work inside the task, not a protocol error. The agent should pause and retry internally, possibly reporting progress in a status update, and if it finally gives up, end the task in TASK_STATE_FAILED with a status message that says why. Do not drop the stream without a terminal event; the client will reconnect and may resubscribe repeatedly.

Push notifications reverse the roles. The agent is now the HTTP client and the caller's webhook can rate limit it. The spec says clients SHOULD rate limit webhooks against flooding, so an agent sending push notifications must honour 429 and Retry-After from webhooks using the same client code, and should coalesce queued status updates for a task rather than replaying each one.

Worked example: an orchestrator and a research agent

An orchestrator fans out work to a research agent whose policy is 60 SendMessage calls per minute per caller, with a burst of 10. The orchestrator submits 30 calls at once. The first 10 are admitted from the burst; the remaining 20 are rejected with 429. Each rejection's RetryInfo gives the time until the next token, one second per token at 60 a minute, so the delays run from 1 to 20 seconds. The client keeps the largest delay per agent, schedules the 20 calls at one-second intervals with a little jitter, and all of them are admitted on their first retry. Total time is about 20 seconds and the agent sees no extra load.

Against an agent that returns a bare 500, the client cannot tell a quota from a crash. It backs off exponentially, the 20 calls retry together and are mostly rejected again, and several are abandoned when the retry budget runs out. Same limiter, worse outcome: the difference is the response format.

Operating it

  • Count rejections by reason, caller and method. A sudden rise for one caller is usually a runaway loop; a rise across callers means the limit is too low or the agent is overloaded.
  • Alert on overload, not on quota. 429s for a caller over its quota are the limiter working. 503s mean the agent is failing for everyone.
  • Test through every proxy. Gateways sometimes strip bodies or rewrite 429 to 502. Send a rejection end to end in a staging test and assert the reason and delay arrive.
  • Publish the policy. The spec allows an extended Agent Card to describe rate limits; document them so callers can plan.

Failure modes and trade-offs

ProblemEffectRemedy
Rate limit sent as plain 500Clients cannot tell it from a crash429 or RESOURCE_EXHAUSTED with ErrorInfo
Retry-After as an HTTP dateClock skew makes delays wrongSend seconds
Same fixed delay for everyoneSynchronised retry wavesExact per-caller delay plus client jitter
Limit checked after task creationRetries create duplicate tasksDecide at admission
Custom A2A code in -32010 and upCollides with a future spec version-32603 plus ErrorInfo in your domain
429 on JSON-RPC to strict librariesBody never parsedSend 200 with the error body to those callers

The trade-off is precision against compatibility: emit the precise form, accept the spec's example codes from others, and keep all meaning in the body.

What to do next

  1. List every limit your agent enforces and confirm each one is checked at admission, before a task is created.
  2. Implement one rejection value and an encoder per binding you serve, with tests that compare reason and delay across encodings.
  3. Return 429 or RESOURCE_EXHAUSTED for caller quota and 503 or UNAVAILABLE for overload, with RetryInfo, Retry-After in seconds and ErrorInfo in your own domain.
  4. Update clients to prefer RetryInfo, parse both Retry-After forms, add jitter, retry with the same messageId and respect a total deadline.
  5. Check streaming paths: no limit may fire after the 200; internal limits end in a FAILED task with a reason.
  6. Send a rejection through every gateway in staging and confirm the body survives.
Key takeaway: A2A 1.0 asks agents to rate limit and to return an appropriate error, lists rate limiting under system errors with example codes, and allows retry guidance; it defines no rate-limit error of its own. Fill that gap with one consistent rejection: 429 or RESOURCE_EXHAUSTED for a caller's quota, 503 or UNAVAILABLE for overload, and a body carrying ErrorInfo in your own domain, RetryInfo with the exact delay and QuotaFailure naming the limit. Decide at admission so a retry is always safe, never fire a limit after a stream has opened, and teach clients to wait exactly as long as they are told, with jitter.