A2A defines one data model and three ways to carry it: JSON-RPC over HTTP, a RESTful HTTP+JSON binding and gRPC. Most examples use JSON-RPC, so the gRPC binding is easy to overlook. It is a good fit for agents that live inside a service mesh, that are called by typed code generated from the same schema, or that stream a lot of updates, because it brings Protocol Buffers, HTTP/2 multiplexing and the gRPC tooling many teams already run.

This page explains the binding as defined by A2A 1.0 and shows how to implement and operate it. Names and rules were checked against specification/a2a.proto and the specification document in the a2aproject/A2A repository at the v1.0.0 tag on 2 October 2026; the proto at the tag matched the main branch for everything cited here. For the JSON-RPC equivalent of the same flows see A2A JSON-RPC request and response flow.

Advertisement

One proto, three bindings

A2A over gRPC: one proto, one service, metadata for service parametersAgent CardsupportedInterfaces[]protocolBinding: GRPCprotocolVersion: 1.0client agentgenerated A2AService stubL7 proxyHTTP/2 aware, TLSHTTP/2 + TLSA2AService serverlf.a2a.v1task storetasks, artifacts, historymetadata on every callauthorization: Bearer ... a2a-version: 1.0a2a-extensions: comma-separated URIs traceparentUnary: SendMessage, GetTask, ListTasks, CancelTask, push config CRUD, GetExtendedAgentCardServer streaming: SendStreamingMessage, SubscribeToTask (StreamResponse: task, message, status_update, artifact_update)Errors: google.rpc.Status with ErrorInfo, domain a2a-protocol.org
The gRPC binding is the proto's A2AService served over HTTP/2 with TLS; service parameters travel as metadata.

In A2A 1.0 the Protocol Buffers file is the normative definition of the data model. Its package is lf.a2a.v1 and it defines one service, A2AService. Each RPC also carries a google.api.http annotation, such as post: "/message:send" for SendMessage, and those annotations define the HTTP+JSON binding. So the gRPC binding is the most direct form of the protocol: the same messages, sent as binary protobuf rather than mapped to JSON.

The specification's requirements for the binding are short: gRPC over HTTP/2 with TLS, proto3 serialization, the normative a2a.proto and an implementation of A2AService.

The eleven RPCs

RPCRequestResponseShape
SendMessageSendMessageRequestSendMessageResponse (task or message)unary
SendStreamingMessageSendMessageRequeststream StreamResponseserver streaming
GetTaskGetTaskRequestTaskunary
ListTasksListTasksRequestListTasksResponseunary
CancelTaskCancelTaskRequestTaskunary
SubscribeToTaskSubscribeToTaskRequeststream StreamResponseserver streaming
CreateTaskPushNotificationConfigTaskPushNotificationConfigTaskPushNotificationConfigunary
GetTaskPushNotificationConfigGetTaskPushNotificationConfigRequestTaskPushNotificationConfigunary
ListTaskPushNotificationConfigsListTaskPushNotificationConfigsRequestListTaskPushNotificationConfigsResponseunary
DeleteTaskPushNotificationConfigDeleteTaskPushNotificationConfigRequestgoogle.protobuf.Emptyunary
GetExtendedAgentCardGetExtendedAgentCardRequestAgentCardunary

Two details matter in practice. First, almost every request message has a tenant string field, which carries the tenant that the HTTP binding puts in the path. Copy it from the Agent Card interface entry, which also has a tenant field. Second, SendMessageResponse and StreamResponse are oneof payloads, so code must branch on which field is set rather than assume a Task comes back.

Advertisement

Declaring a gRPC interface in the Agent Card

A 1.0 Agent Card lists its endpoints in supportedInterfaces, in order of preference. Each AgentInterface has a url, a protocolBinding whose core values are JSONRPC, GRPC and HTTP+JSON, an optional tenant and a protocolVersion such as 1.0. The rest of the card is covered in the Agent Card specification, field by field.

"supportedInterfaces": [
  {"url": "https://grpc.agents.example.com", "protocolBinding": "GRPC",     "protocolVersion": "1.0"},
  {"url": "https://agents.example.com/recon/a2a", "protocolBinding": "JSONRPC", "protocolVersion": "1.0"}
]

There is one ambiguity to handle deliberately. The proto comment says the URL must be an absolute HTTPS URL in production and gives https://grpc.example.com/a2a as an example, but standard gRPC clients connect to a host and port and send every call to a fixed path made of the service and method name, such as /lf.a2a.v1.A2AService/SendMessage. A path segment in the URL has no place in ordinary gRPC addressing. The safest choice is a gRPC URL with no path. If you must publish one with a path, put a proxy in front that strips it, and test the SDKs your callers use, because they may not agree on how to treat it.

Service parameters travel as metadata

A2A defines service parameters that are not part of any message, chiefly the protocol version and the extensions a client wants. The gRPC binding requires them to be sent as metadata. Keys are case-insensitive and gRPC lowercases them, so they appear as a2a-version and a2a-extensions; several extension URIs go in one comma-separated value. Credentials go in authorization like any gRPC call.

The version deserves attention. The specification says an empty version must be interpreted as 0.3, so a client that forgets the header talks to a 1.0 server as if it were a 0.3 client. Always send a2a-version: 1.0. On the server, read the value and return VersionNotSupported for a version the interface does not serve, rather than guessing.

A client in Python

Generate stubs from the proto with grpcio-tools. The proto imports Google API annotations, so the googleapis protos must be on the include path at generation time, and googleapis-common-protos must be installed at run time.

python -m grpc_tools.protoc -I specification -I third_party/googleapis \
  --python_out=gen --grpc_python_out=gen specification/a2a.proto
import uuid, grpc
import a2a_pb2 as pb, a2a_pb2_grpc as rpc

META = [("authorization", "Bearer " + TOKEN), ("a2a-version", "1.0")]
channel = grpc.secure_channel("grpc.agents.example.com:443", grpc.ssl_channel_credentials())
stub = rpc.A2AServiceStub(channel)

req = pb.SendMessageRequest(message=pb.Message(
    message_id=str(uuid.uuid4()), role=pb.ROLE_USER,
    parts=[pb.Part(text="Reconcile invoice INV-1042 against PO-77")]))

task_id = None
for ev in stub.SendStreamingMessage(req, metadata=META, timeout=300):
    kind = ev.WhichOneof("payload")
    if kind == "task":
        task_id = ev.task.id
    elif kind == "message":                     # agent answered directly, no task
        print(ev.message.parts[0].text)
    elif kind == "status_update":
        print("state:", pb.TaskState.Name(ev.status_update.status.state))
    elif kind == "artifact_update":
        a = ev.artifact_update
        print("artifact", a.artifact.artifact_id, "append" if a.append else "new", a.last_chunk)
# Stream ended: confirm the outcome rather than inferring it
if task_id:
    print(stub.GetTask(pb.GetTaskRequest(id=task_id), metadata=META, timeout=10).status.state)

A server, including rich errors

On the server, implement the generated servicer. Streaming RPCs are generators that yield StreamResponse messages. A2A-specific errors must carry a google.rpc.ErrorInfo in the status details, with the reason in upper snake case without the Error suffix and the domain a2a-protocol.org. The grpcio-status package converts a google.rpc.Status into something the server can abort with.

from google.protobuf import any_pb2
from google.rpc import code_pb2, error_details_pb2, status_pb2
from grpc_status import rpc_status

def a2a_abort(context, code, reason, message, **meta):
    info = error_details_pb2.ErrorInfo(reason=reason, domain="a2a-protocol.org", metadata=meta)
    detail = any_pb2.Any(); detail.Pack(info)
    context.abort_with_status(rpc_status.to_status(
        status_pb2.Status(code=code, message=message, details=[detail])))

class Agent(rpc.A2AServiceServicer):
    def SubscribeToTask(self, request, context):
        md = dict(context.invocation_metadata())
        if md.get("a2a-version", "") != "1.0":
            a2a_abort(context, code_pb2.FAILED_PRECONDITION, "VERSION_NOT_SUPPORTED", "Send a2a-version: 1.0")
        task = STORE.get(request.id, principal_of(context))
        if task is None:
            a2a_abort(context, code_pb2.NOT_FOUND, "TASK_NOT_FOUND", "Task not found", taskId=request.id)
        if is_terminal(task.status.state):
            a2a_abort(context, code_pb2.FAILED_PRECONDITION, "UNSUPPORTED_OPERATION", "Task already finished")
        yield pb.StreamResponse(task=task)
        for event in STORE.events_after(task.id):     # blocks until new events
            yield event
            if closes_stream(event):
                return

The status codes are coarse: TaskNotFound maps to NOT_FOUND, but TaskNotCancelable, UnsupportedOperation, VersionNotSupported and several others all map to FAILED_PRECONDITION. Clients must branch on the ErrorInfo reason. The full table and a client classifier that decides what to retry are in A2A error codes, in depth.

Streaming semantics

A stream either carries exactly one Message and closes, or begins with the Task and then carries status and artifact update events. The specification requires the stream to close when the task reaches a terminal state (completed, failed, canceled or rejected), and its streaming section also says the stream closes at an interrupted state, meaning input-required or auth-required. Write clients that accept either: when a stream ends, call GetTask and act on the state you read, and call SubscribeToTask to resume a task that is still running. SubscribeToTask on a terminal task returns UnsupportedOperation, so a client that resubscribes blindly will see that error at the end of every task.

Network drops look like a stream ending with an UNAVAILABLE or CANCELLED status rather than a clean close. Treat them the same way: read the task, then resubscribe. Artifact chunks may arrive again after a reconnect, so assemble artifacts by artifact ID and apply append idempotently. The cross-binding view of the same lifecycle is in A2A streaming architecture.

A worked trace with a dropped connection

Follow the invoice reconciliation request from the client example through one realistic run. The numbers are illustrative; the sequence is what the binding prescribes.

  1. The client opens SendStreamingMessage with a2a-version: 1.0 and a 300-second deadline. The first StreamResponse carries the Task, with ID t-91 and state TASK_STATE_SUBMITTED. The client records the ID before doing anything else, because it is the only handle for recovery.
  2. A status_update moves the task to TASK_STATE_WORKING. Two artifact_update events follow for artifact recon-report: the first with append false, the second with append true and last_chunk false.
  3. Forty seconds pass without events while the agent queries a ledger system. A proxy with a 30-second idle timeout closes the connection, and the client sees the stream fail with UNAVAILABLE. Nothing is wrong with the task; only the transport ended.
  4. The client calls GetTask for t-91. The task is still TASK_STATE_WORKING and its artifacts list already holds the first two chunks, so the client's local copy is consistent.
  5. The client calls SubscribeToTask. The stream starts with the current Task again, then delivers the final chunk with last_chunk true and a status_update to TASK_STATE_COMPLETED, and the server closes the stream.
  6. Because the first event after resubscribing is a full Task snapshot, the client replaces its local state from it instead of appending, which removes any risk of duplicating the chunks it already had.

Two fixes come out of this trace. Raise the proxy idle timeout and enable keepalive pings so step 3 does not happen in normal operation, and keep the recovery path anyway, because deploys, node failures and mobile networks will cut streams regardless.

Running it behind real infrastructure

  • Load balancing: gRPC multiplexes many calls on one long-lived HTTP/2 connection, so a layer-4 balancer pins all of a client's traffic to one backend. Use an HTTP/2-aware layer-7 proxy or client-side balancing.
  • Idle streams: a task that is quiet for minutes keeps its stream open. Configure HTTP/2 keepalive pings on both ends and raise proxy idle timeouts above your longest quiet period, or streams will be cut and clients will resubscribe constantly.
  • Deadlines: set a deadline on every unary call and a generous one on streams, and propagate the remaining budget when an agent delegates further; see A2A timeout handling.
  • Message size: gRPC implementations commonly default to a 4 MB maximum received message. Inline raw file parts can exceed it; prefer url parts for large files rather than raising the limit everywhere.
  • Tracing and health: pass traceparent in metadata so traces cross the hop, and expose the standard gRPC health service for probes; it is not part of A2A but every gRPC load balancer understands it.

Failure modes and trade-offs

  • Silent version downgrade: a missing a2a-version is read as 0.3. Reject requests without it in 1.0-only deployments.
  • Path in the card URL: clients connect to the host and ignore or mishandle the path. Publish path-less gRPC URLs.
  • Status-code-only error handling: treating every FAILED_PRECONDITION alike retries version mismatches forever. Read ErrorInfo.
  • Schema drift: stubs generated from an older proto keep unknown fields as opaque bytes but cannot read them, so new data is invisible to old code. Pin the proto version and regenerate deliberately.
  • Browser callers: browsers cannot speak native gRPC, so serve JSON-RPC or HTTP+JSON as well when web clients need the agent.
  • The trade-off: gRPC gives typed contracts, compact messages and efficient streaming, and costs an HTTP/2-aware proxy layer and harder debugging with ordinary HTTP tools. Offer it alongside a JSON binding, and list it first in the card only when most of your callers are services that already use gRPC.

What to do next

  1. Generate stubs from the v1.0.0 a2a.proto and pin that version in your build.
  2. Add a GRPC entry to your Agent Card's supportedInterfaces with a path-less URL and protocolVersion 1.0.
  3. Send a2a-version: 1.0 on every client call and reject missing or unsupported versions on the server.
  4. Implement errors as google.rpc.Status with ErrorInfo and test that clients branch on the reason, not the code.
  5. Test stream recovery by killing the connection mid-task and checking that GetTask plus SubscribeToTask resumes cleanly.
  6. Configure keepalive, proxy idle timeouts, deadlines and message size limits, and run the gRPC health service.
Key takeaway: The A2A gRPC binding serves the normative lf.a2a.v1 A2AService over HTTP/2 with TLS, with eleven RPCs, two of them server-streaming. Declare it as a GRPC interface in the Agent Card with a path-less URL, always send a2a-version 1.0 in metadata, return google.rpc.Status with an A2A ErrorInfo and branch on its reason. Treat every stream end as a cue to read the task and resubscribe, and put an HTTP/2-aware proxy, keepalives and deadlines in front before production traffic arrives.