An agent built with ADK for Java is a library object. LlmAgent holds the model, instruction and tools, and a Runner executes a turn and emits an RxJava Flowable<Event>. Nothing in that picture says how another service reaches the agent. Teams that already run gRPC between Java services usually want the agent to be one more gRPC service: a typed contract, generated clients in every language, HTTP/2 streaming, deadlines and cancellation that propagate, and the same load balancers and mesh policies as everything else.

This article builds that front end. It defines a protobuf contract for sessions and runs, bridges the event stream to a gRPC server stream without losing flow control or cancellation, configures the server for long-lived streaming calls, and covers errors, retries, load balancing and the main alternative: the gRPC binding of the A2A protocol. The ADK calls used are the ones in the official Java quickstart (com.google.adk:google-adk), plus the google-genai Content and Part types for messages; everything else is plain grpc-java.

Two ways to put gRPC in front of an agent

There are two reasonable gRPC shapes, and they solve different problems; REST is listed as the non-gRPC baseline for comparison.

OptionContractBest for
Your own proto around the Runneryou design it: sessions, runs, event fieldsinternal callers you control, polyglot clients, strict typing
A2A over its gRPC bindingthe A2A spec's service and messagesagents from other teams or vendors that discover each other through agent cards
REST or server-sent eventsJSON over HTTP/1.1browsers and simple integrations

A private proto is simplest when you own both ends: it can expose exactly the event fields your callers need. A2A buys interoperability at the cost of a larger, versioned surface. Both put the same bridging problem in the middle, so we build the private contract first.

The contract

Keep the contract small and explicit about identity. A run needs a user, a session and a message; the response is a stream because an agent turn can produce several events (tool calls, tool results, partial text, a final answer). The request_id field exists for idempotency, discussed below.

syntax = "proto3";
package agents.v1;
option java_multiple_files = true;
option java_package = "com.example.agents.v1";

service AgentService {
  rpc CreateSession(CreateSessionRequest) returns (SessionRef);
  rpc Run(RunRequest) returns (stream AgentEvent);
}

message CreateSessionRequest { string user_id = 1; }
message SessionRef { string user_id = 1; string session_id = 2; }

message RunRequest {
  string user_id = 1;
  string session_id = 2;
  string text = 3;
  string request_id = 4;   // caller-generated, unique per logical turn
}

message AgentEvent {
  int32 seq = 1;           // 0, 1, 2 ... within one Run
  string text = 2;         // rendered event content
  bool is_final = 3;       // true on the agent's final response
}

Generate Java stubs with the protobuf Maven or Gradle plugin and depend on grpc-netty-shaded, grpc-protobuf, grpc-stub and grpc-services at matching versions. Version the package (agents.v1) from day one; adding fields is compatible, renaming or renumbering them is not.

The agent and runner

The agent and runner come straight from the quickstart pattern. InMemoryRunner is fine for a single replica and for tests; with more than one replica, construct a Runner over a persistent session service so any replica can continue any session (see ADK Java runtime, session and context).

LlmAgent agent = LlmAgent.builder()
    .name("order_helper")
    .description("Answers questions about customer orders")
    .instruction("Use lookupOrder before answering any order question.")
    .model("gemini-flash-latest")
    .tools(FunctionTool.create(OrderTools.class, "lookupOrder"))
    .build();

InMemoryRunner runner = new InMemoryRunner(agent);

Bridging Flowable events to a gRPC stream

This is the heart of the integration. Three things must hold. Events reach the client in order. A slow client must not make the server buffer without limit. A client that cancels or hits its deadline must stop the model call and its tools, not just stop reading.

grpc-java exposes all three through ServerCallStreamObserver: isReady() and setOnReadyHandler for transport flow control, and setOnCancelHandler for cancellation. The handlers must be registered before the service method returns, so register them first and let them find the subscription through a reference. On the RxJava side, request events one at a time while the transport is ready.

public final class AgentGrpcService extends AgentServiceGrpc.AgentServiceImplBase {
  private final Runner runner;
  AgentGrpcService(Runner runner) { this.runner = runner; }

  @Override
  public void createSession(CreateSessionRequest req, StreamObserver<SessionRef> out) {
    runner.sessionService().createSession(runner.appName(), req.getUserId())
        .subscribe(s -> {
              out.onNext(SessionRef.newBuilder()
                  .setUserId(s.userId()).setSessionId(s.id()).build());
              out.onCompleted();
            },
            err -> out.onError(toStatus(err).asRuntimeException()));
  }

  @Override
  public void run(RunRequest req, StreamObserver<AgentEvent> obs) {
    var out = (ServerCallStreamObserver<AgentEvent>) obs;
    var upstream = new AtomicReference<Subscription>();
    var seq = new AtomicInteger();
    out.setOnCancelHandler(() -> { var s = upstream.get(); if (s != null) s.cancel(); });
    out.setOnReadyHandler(() -> { var s = upstream.get(); if (s != null && out.isReady()) s.request(1); });

    Content msg = Content.fromParts(Part.fromText(req.getText()));
    runner.runAsync(req.getUserId(), req.getSessionId(), msg, RunConfig.builder().build())
        .onBackpressureBuffer(256)                 // bounded; overflow becomes an error
        .subscribe(new FlowableSubscriber<Event>() {
          public void onSubscribe(Subscription s) { upstream.set(s); s.request(1); }
          public void onNext(Event e) {
            if (out.isCancelled()) return;
            out.onNext(AgentEvent.newBuilder().setSeq(seq.getAndIncrement())
                .setText(e.stringifyContent()).setIsFinal(e.finalResponse()).build());
            if (out.isReady()) upstream.get().request(1);
          }
          public void onError(Throwable t) { out.onError(toStatus(t).asRuntimeException()); }
          public void onComplete() { out.onCompleted(); }
        });
  }
}

Requests are additive, so the ready handler and onNext occasionally both requesting is harmless: at worst one or two extra events are in flight. Whether the upstream truly slows down depends on how the runner produces events. The model streams at its own pace, so the onBackpressureBuffer(256) cap is what bounds memory. A client that falls 256 events behind gets an error instead of exhausting the heap, which is a policy choice you should make on purpose. Reactive Streams serialises onNext, onError and onComplete, which matters because a gRPC StreamObserver is not thread-safe. For how ADK events are structured and when partial events appear, see ADK Java streaming events.

Server settings for long streaming calls

A streaming agent call can last a minute or more, far longer than a typical unary RPC, and the server settings have to reflect that.

HealthStatusManager health = new HealthStatusManager();
Server server = NettyServerBuilder.forPort(8443)
    .addService(new AgentGrpcService(runner))
    .addService(health.getHealthService())
    .keepAliveTime(60, TimeUnit.SECONDS)          // detect dead peers on idle streams
    .permitKeepAliveTime(30, TimeUnit.SECONDS)    // tolerate client pings this often
    .maxConnectionAge(30, TimeUnit.MINUTES)       // force periodic reconnects for rebalancing
    .maxConnectionAgeGrace(10, TimeUnit.MINUTES)  // longer than your longest run
    .maxInboundMessageSize(4 * 1024 * 1024)
    .build()
    .start();
health.setStatus("", HealthCheckResponse.ServingStatus.SERVING);

Runtime.getRuntime().addShutdownHook(new Thread(() -> {
  health.enterTerminalState();                    // stop new traffic first
  server.shutdown();                              // let in-flight runs finish
  try { server.awaitTermination(10, TimeUnit.MINUTES); }
  catch (InterruptedException e) { server.shutdownNow(); }
}));

The connection-age grace and the shutdown wait must both exceed your longest expected run; otherwise rollouts cut answers off mid-sentence. Keepalive values must agree with the client and with any proxy in between: a client pinging more often than permitKeepAliveTime allows gets its connection closed with a GOAWAY.

The client, deadlines and a worked trace

Every call carries a deadline. Without one, a stuck tool holds the stream, a server thread and a model request open indefinitely.

ManagedChannel channel = ManagedChannelBuilder.forTarget("dns:///agents.internal:8443")
    .defaultLoadBalancingPolicy("round_robin")
    .useTransportSecurity()
    .build();
var stub = AgentServiceGrpc.newBlockingStub(channel);

var it = stub.withDeadlineAfter(120, TimeUnit.SECONDS).run(RunRequest.newBuilder()
    .setUserId("u-42").setSessionId(sessionId)
    .setText("Where is order 1187?")
    .setRequestId(UUID.randomUUID().toString()).build());
while (it.hasNext()) {
  AgentEvent e = it.next();
  if (e.getIsFinal()) System.out.println(e.getText());
}

The deadline travels to the server, where it cancels the call context, fires the cancel handler and cancels the RxJava subscription. In a worked trace for the order question, the client sees event 0 (the model's call to lookupOrder), event 1 (the tool's result), and event 2 (the final answer, is_final true), then the stream completes with status OK.

One Run call: a unary request in, a server stream of agent events outCallerstub + deadlinegRPC serverNetty, HTTP/2RunRequeststream AgentEventAgentGrpcServiceflow-control bridgeADK RunnerrunAsyncrequest(1)EventLlmAgentmodel + toolsSessionServicestate, historyHealth serviceSERVING / NOTCancel / deadlineclient gives upcancels subscriptionCancellation flows right to left and stops the model call; readiness flows left to right and paces events.
The bridge passes events forward only while the transport is ready, and turns client cancellation into cancellation of the agent run.

Errors, status codes and retries

Map failures to status codes deliberately; callers' retry logic keys off them.

SituationStatusCaller should
Missing user, session or textINVALID_ARGUMENTfix the request; never retry
Unknown sessionNOT_FOUNDcreate a session, then retry once
Client fell behind the buffer capRESOURCE_EXHAUSTEDread faster or use a smaller turn
Run exceeded the deadlineDEADLINE_EXCEEDEDretry only if the turn is safe to repeat
Model or tool failureUNAVAILABLE or INTERNALback off; surface a safe message

Keep toStatus a small, tested function. Put a sanitised description in the status and the full exception in your logs. Model errors can echo prompt text, and status descriptions travel to callers.

Retries need care. An agent turn is not idempotent: tools may have sent an email or issued a refund before the stream broke. Do not put Run in a service-config retry policy. Instead, record request_id with the turn's outcome in the session or a side table. If a repeated request arrives, replay the stored result or report that the turn is in progress rather than running tools twice.

Testing the bridge in-process

Test the bridge in-process, without a network or a real model. grpc-java's in-process transport runs the real server and client stacks in one JVM, so flow control, deadlines and cancellation behave as they do in production. Back the runner with a stub agent or tool that sleeps, then assert three behaviours: events arrive in seq order, a short deadline ends the call with DEADLINE_EXCEEDED and the tool observes cancellation, and a reader that stalls produces RESOURCE_EXHAUSTED rather than unbounded memory growth.

String name = InProcessServerBuilder.generateName();
Server server = InProcessServerBuilder.forName(name).directExecutor()
    .addService(new AgentGrpcService(testRunner)).build().start();
ManagedChannel ch = InProcessChannelBuilder.forName(name).directExecutor().build();

var stub = AgentServiceGrpc.newBlockingStub(ch).withDeadlineAfter(200, TimeUnit.MILLISECONDS);
var ex = assertThrows(StatusRuntimeException.class,
    () -> stub.run(slowRequest()).forEachRemaining(e -> {}));
assertEquals(Status.Code.DEADLINE_EXCEEDED, ex.getStatus().getCode());
assertTrue(slowTool.wasCancelled());

Load balancing, health and observability

gRPC multiplexes calls over long-lived HTTP/2 connections, which defeats connection-level (L4) load balancing: one busy client can pin all its streams to one replica for the life of the connection. Use client-side round_robin against a headless service, or an L7 proxy or mesh that balances per request. Keep maxConnectionAge set so new replicas get traffic after a scale-out. Kubernetes readiness can use the standard gRPC health protocol directly; deploying ADK Java on Kubernetes covers the rest of the manifest.

For observability (the broader setup is in ADK Java observability), add a server interceptor that records method, status code, stream duration and events-per-run. Propagate trace context through gRPC metadata so a slow tool call shows up inside the caller's trace. Alert on a rising RESOURCE_EXHAUSTED or CANCELLED rate: the first means slow consumers, the second usually means deadlines set shorter than real runs.

The alternative: A2A over gRPC

If the caller is another team's agent rather than your own service, consider the Agent2Agent (A2A) protocol, whose specification includes a gRPC binding alongside JSON-RPC and HTTP+JSON. In the current specification proto the service is A2AService with RPCs including SendMessage, SendStreamingMessage, GetTask, CancelTask and SubscribeToTask. The protocol has evolved, and RPC and field names changed between versions, so pin one spec version and generate code from that version's proto rather than copying names from blog posts.

The A2A Java SDK (io.github.a2asdk) provides a gRPC reference server module, a2a-java-sdk-reference-grpc, and a client transport, a2a-java-sdk-client-transport-grpc. An agent card advertises which transports an agent serves. The bridging problem does not go away: you still adapt the ADK runner's events to A2A task and message updates. See ADK Java and A2A integration for the protocol model.

Failure modes

  • Unbounded buffering. Ignoring isReady() lets a slow mobile client grow the server heap until it dies.
  • Orphaned runs. Without a cancel handler, a client that disconnects leaves the model and tools running and billing.
  • Rollouts that truncate answers because the connection-age grace or shutdown wait is shorter than a run.
  • Duplicate side effects from automatic retries of a non-idempotent Run.
  • Lost sessions when InMemoryRunner runs behind a balancer with several replicas.

Trade-offs

gRPC gives typed contracts, efficient streaming and first-class deadlines, at the cost of browser support (which needs gRPC-Web or a gateway), harder debugging than JSON over HTTP, and load-balancing work. For internal service-to-agent traffic it is usually the right default. For public or browser clients, put REST with server-sent events at the edge and keep gRPC inside. For cross-organisation agents, A2A's interoperability is worth its extra surface.

What to do next

  1. Write the agents.v1 proto and generate stubs; review field names as an API, not an implementation detail.
  2. Implement the bridge with ready and cancel handlers, and test it with a client that reads slowly and one that cancels mid-run.
  3. Configure keepalive, connection age and shutdown waits from your measured longest run.
  4. Add request_id deduplication before enabling any retries.
  5. Replace InMemoryRunner with a persistent session service before running a second replica.
  6. Add an interceptor for status, duration and events per run, and alert on RESOURCE_EXHAUSTED and CANCELLED.
Key takeaway: ADK for Java gives you an agent, a Runner and a Flowable of events; gRPC is a front end you add. Define a small versioned proto, bridge the event stream with isReady and setOnReadyHandler for flow control and setOnCancelHandler for cancellation, give every call a deadline, size keepalive and shutdown settings to your longest run, never auto-retry a non-idempotent turn, and balance per request rather than per connection. Use the A2A gRPC binding when the caller is another organisation's agent.