Kubernetes was tuned for web services whose requests take milliseconds and whose cost is CPU. An agent built with the Agent Development Kit for Java looks different. A single turn can run for a minute or more while the model streams tokens and tools are called, the process spends nearly all of that time waiting on the network, and every replica draws on the same model API quota. The default recipe of a Deployment, a CPU-based autoscaler and an HTTP health check deploys such a service fine and then misbehaves under load and during every rollout.
This article covers the Kubernetes side: the image, the probes, the arithmetic behind the termination grace period, which metric to autoscale on, how to protect capacity during node maintenance, and how to give pods model credentials without key files. Packaging choices and the alternatives to Kubernetes, such as Cloud Run, are covered in ADK Java deployment, and the in-process drain logic in graceful shutdown for ADK Java. Here we connect them to the cluster.
What makes agent traffic different
Three properties drive every decision below. First, turns are long. A request that calls the model three times with two tool calls in between can take 30 to 120 seconds, and with streaming the HTTP response is open the whole time. Anything that cuts connections, such as a rollout, a node drain or a load balancer idle timeout, cuts conversations in the middle.
Second, the work is I/O-bound. The JVM thread handling a turn is almost always blocked on a socket, so CPU utilisation stays low even when the pod is saturated with concurrent conversations. A CPU-based autoscaler never sees the load.
Third, the expensive resource is outside the cluster. Model calls are limited by a per-project quota, so beyond a point more replicas just convert queueing into HTTP 429 errors.
State must also live outside the pod. The Runner appends every non-partial event to its session service, and InMemoryRunner keeps sessions in process memory, so with more than one replica the next turn of a conversation can land on a pod that has never seen it. Use a durable BaseSessionService implementation; a relational design is described in the ADK Postgres session table schema.
The container image and JVM sizing
Build a small, non-root image with no shell. The runtime image below is Google's distroless Java 21 image; any maintained JRE base works. The ADK dependency is com.google.adk:google-adk; pin the version you tested, because the project moves quickly.
# Multi-stage build; the runtime image has no shell and runs as non-root.
FROM eclipse-temurin:21-jdk AS build
WORKDIR /src
COPY . .
RUN ./mvnw -q -DskipTests package
FROM gcr.io/distroless/java21-debian12:nonroot
WORKDIR /app
COPY --from=build /src/target/agent-service.jar /app/agent-service.jar
# Size the heap from the container limit, leave room for metaspace, threads and direct buffers.
ENV JAVA_TOOL_OPTIONS="-XX:MaxRAMPercentage=70 -XX:+ExitOnOutOfMemoryError"
EXPOSE 8080
ENTRYPOINT ["java", "-jar", "/app/agent-service.jar"]Modern JVMs read the container's memory limit, so size the heap as a percentage of it rather than with a fixed -Xmx. Keep about 30 percent for non-heap memory: metaspace, thread stacks, the gRPC or HTTP client's direct buffers and the JIT. -XX:+ExitOnOutOfMemoryError turns a heap exhaustion into a clean restart instead of a half-alive process. Set the memory limit equal to the request so the pod is never killed for using memory the scheduler promised it, and leave the CPU limit off, or set it high, because CPU throttling during JIT warm-up makes startup slow and erratic.
Probes that mean something
A pod should receive traffic only when it can actually complete a turn. That means the model name resolves, credentials work, the session store answers and egress to the model endpoint is open. Spring Boot's availability support gives you the plumbing. With probes enabled it exposes separate liveness and readiness groups, and it publishes the accepting-traffic readiness state only after every ApplicationRunner has completed. Put the warm-up in a runner and readiness is gated on it for free:
# application.yaml
server:
shutdown: graceful # set explicitly: stop accepting, let in-flight requests finish
spring:
threads:
virtual:
enabled: true # request threads block on model calls cheaply (Java 21+)
lifecycle:
timeout-per-shutdown-phase: 90s # must fit inside terminationGracePeriodSeconds
management:
endpoint:
health:
probes:
enabled: true # /actuator/health/liveness and /actuator/health/readiness
endpoints:
web:
exposure:
include: health,prometheus@Component
class AgentWarmup implements ApplicationRunner {
private final Runner runner;
private final AgentProperties props;
AgentWarmup(Runner runner, AgentProperties props) { this.runner = runner; this.props = props; }
// Spring Boot publishes ReadinessState.ACCEPTING_TRAFFIC only after runners complete,
// so a failure here keeps the pod unready and halts a bad rollout.
@Override public void run(ApplicationArguments args) {
LlmRegistry.getLlm(props.model()); // fail fast on a bad model name
Session s = runner.sessionService()
.createSession(runner.appName(), "warmup").blockingGet();
runner.runAsync("warmup", s.id(), Content.fromParts(Part.fromText("Reply with: ready")))
.timeout(30, TimeUnit.SECONDS)
.blockingSubscribe();
}
}
@Service
class AgentService {
private final Runner runner;
private final AtomicInteger inFlight = new AtomicInteger();
AgentService(Runner runner, MeterRegistry meters) {
this.runner = runner;
meters.gauge("adk_inflight_runs", inFlight); // the autoscaling signal
}
Flowable<Event> turn(String user, String sessionId, String text) {
return runner.runAsync(user, sessionId, Content.fromParts(Part.fromText(text)))
.doOnSubscribe(s -> inFlight.incrementAndGet())
.doFinally(inFlight::decrementAndGet);
}
}The warm-up follows the boot pattern in the ADK Java runtime boot sequence: ADK resolves the model lazily, so an explicit LlmRegistry.getLlm call and one real run move configuration errors to startup. If the warm-up throws, the application fails to start, the new replica never becomes ready, and with maxUnavailable: 0 the rollout stalls while the old version keeps serving.
Keep the model out of the recurring probes. Liveness should answer only the question of whether this JVM is wedged. If liveness or readiness called the model, a provider outage would mark every replica unready at once, or restart all of them in a loop, turning a partial degradation into a total one.
The Deployment, and the grace-period arithmetic
apiVersion: apps/v1
kind: Deployment
metadata: {name: support-agent}
spec:
replicas: 3
strategy: {type: RollingUpdate, rollingUpdate: {maxSurge: 1, maxUnavailable: 0}}
selector: {matchLabels: {app: support-agent}}
template:
metadata: {labels: {app: support-agent}}
spec:
serviceAccountName: support-agent # bound to a cloud identity, no key file
terminationGracePeriodSeconds: 120 # preStop 10s + drain 90s + margin
topologySpreadConstraints:
- {maxSkew: 1, topologyKey: topology.kubernetes.io/zone,
whenUnsatisfiable: ScheduleAnyway, labelSelector: {matchLabels: {app: support-agent}}}
containers:
- name: agent
image: registry.example.com/support-agent:1.4.2
ports: [{containerPort: 8080}]
env:
- {name: AGENT_MODEL, value: gemini-2.5-flash}
- {name: GOOGLE_CLOUD_PROJECT, value: my-project}
- {name: GOOGLE_CLOUD_LOCATION, value: us-central1}
- name: SESSION_DB_PASSWORD
valueFrom: {secretKeyRef: {name: session-db, key: password}}
resources:
requests: {cpu: "500m", memory: 1Gi}
limits: {memory: 1Gi} # memory limit = request; no CPU limit
startupProbe:
httpGet: {path: /actuator/health/readiness, port: 8080}
periodSeconds: 5
failureThreshold: 24 # up to 2 minutes for JVM start + warm-up
readinessProbe:
httpGet: {path: /actuator/health/readiness, port: 8080}
periodSeconds: 5
livenessProbe:
httpGet: {path: /actuator/health/liveness, port: 8080}
periodSeconds: 10
failureThreshold: 6
lifecycle:
preStop: {sleep: {seconds: 10}} # let endpoint removal propagate firstWhen a pod is deleted, two things start at the same moment: the kubelet runs the preStop hook and then sends SIGTERM, and the control plane removes the pod from Service endpoints. Endpoint removal takes a few seconds to reach every load balancer and kube-proxy, so without a pause the pod stops accepting connections while traffic is still being sent to it. The preStop sleep action covers that gap without needing a shell in the image; it is beta and enabled by default since Kubernetes v1.30. On older clusters, sleep inside the application's shutdown hook instead.
After SIGTERM, Spring's graceful shutdown stops accepting new requests and waits up to timeout-per-shutdown-phase for in-flight ones. The budget is therefore: preStop 10 s, plus the longest turn you are willing to let finish, 90 s here, plus a margin for closing clients and flushing telemetry. At 120 s the kubelet sends SIGKILL regardless. Measure your p99 turn duration before choosing these numbers; if some turns legitimately run for ten minutes, they need checkpointing and resumption, not a ten-minute grace period that slows every rollout.
Check the timeouts on everything in front of the pod too. An ingress or load balancer idle timeout shorter than a long streaming turn will cut the connection even when Kubernetes behaves perfectly.
Autoscaling on the right signal
The metric that tracks load for an I/O-bound agent is the number of turns in flight per pod. The adk_inflight_runs gauge above counts subscriptions to runAsync that have not finished. Expose it through Prometheus, make it available to the autoscaler with a custom-metrics adapter such as the Prometheus adapter or KEDA, and target an average per pod:
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata: {name: support-agent}
spec:
scaleTargetRef: {apiVersion: apps/v1, kind: Deployment, name: support-agent}
minReplicas: 3
maxReplicas: 4 # quota supports 3 pods, plus 1 for failover
metrics:
- type: Pods
pods:
metric: {name: adk_inflight_runs}
target: {type: AverageValue, averageValue: "40"}
behavior:
scaleDown:
stabilizationWindowSeconds: 600 # turns are long; avoid flapping
policies: [{type: Pods, value: 1, periodSeconds: 120}]
---
apiVersion: policy/v1
kind: PodDisruptionBudget
metadata: {name: support-agent}
spec:
minAvailable: 2
selector: {matchLabels: {app: support-agent}}Pick the target by load test. Raise concurrency on one pod until p95 time to first token or memory degrades, then take about 70 percent of that. With virtual threads, covered in virtual threads in ADK Java, a blocked turn costs a small heap object rather than a platform thread, so the limit is usually memory for session histories and response buffers, not threads.
Then cap the top from the quota side. A worked example: the project's model quota is 600 requests per minute, and an average turn makes 3 model calls and lasts 30 seconds, so each in-flight turn consumes about 6 requests per minute. The quota therefore sustains about 100 concurrent turns. At a target of 40 per pod, three replicas already cover that, and twelve would let the service accept 480 concurrent turns when the quota serves about 100, which is why the manifest stops at four. Set maxReplicas to what the quota supports plus headroom for failover, and when the ceiling is reached, reject new turns quickly with a clear retry message rather than queueing them into timeouts.
Disruptions, rollouts and topology
Node upgrades and cluster autoscaler scale-downs evict pods through the eviction API, which respects PodDisruptionBudgets. With minAvailable: 2, voluntary evictions proceed only while at least two replicas stay ready, so maintenance never takes the service to zero. A PDB does not protect against node crashes, which is what the zone topology spread constraint is for.
For rollouts, maxUnavailable: 0 with maxSurge: 1 means a new pod must pass its warm-up before an old one is terminated, and each old pod gets its full drain. Rollouts are slower, but no conversation is cut and a broken build never replaces a working one.
Credentials, secrets and egress
The model client needs credentials. The Java GenAI client that ADK uses reads GOOGLE_API_KEY for the Gemini Developer API, and GOOGLE_CLOUD_PROJECT and GOOGLE_CLOUD_LOCATION plus a toggle variable for the Google Cloud path. The toggle has been renamed across releases: older material uses GOOGLE_GENAI_USE_VERTEXAI and the current client README shows GOOGLE_GENAI_USE_ENTERPRISE, so check the documentation for the client version on your classpath.
On GKE, prefer Workload Identity Federation for GKE over API keys: bind the pod's Kubernetes service account to an IAM identity with permission to call the model, and the client obtains short-lived tokens from the metadata server. There is no key file to leak or rotate. If you must use an API key, mount it from a Secret, never bake it into the image or a ConfigMap, and restrict who can read Secrets in the namespace.
Add a NetworkPolicy that denies egress by default and allows DNS, the model endpoint, the session store and each tool backend. An agent that can be steered by prompt injection into calling arbitrary URLs is far less dangerous when the network refuses to route them.
Failure modes
| Symptom | Cause | Fix |
|---|---|---|
| Conversations cut during every deploy | no preStop pause, or grace period shorter than turns | preStop sleep, grace = pause + p99 turn + margin |
| Pods saturated, HPA does nothing | scaling on CPU for an I/O-bound service | scale on in-flight runs |
| Scale-out makes errors worse | replicas exceed model quota; 429s | derive maxReplicas from quota, shed load early |
| Follow-up turn forgets the conversation | InMemoryRunner with several replicas | durable session service |
| All replicas restart during a provider outage | liveness probe calls the model | model only in startup warm-up |
| OOMKilled under load | heap sized near the limit, no room for buffers | MaxRAMPercentage about 70, limit equals request |
What to do next
- Measure p50, p95 and p99 turn duration and model calls per turn from your traces; every number below depends on them.
- Replace InMemoryRunner with a durable session service and prove that a conversation survives a pod deletion.
- Add the warm-up ApplicationRunner and the three probes, then deploy a bad model name to staging and confirm the rollout stalls.
- Set the preStop sleep, graceful shutdown and grace period from the measured p99, and run a rollout during a load test with zero cut streams as the pass condition.
- Export the in-flight gauge, load-test one pod to find its limit, and configure the HPA target at about 70 percent of it.
- Compute maxReplicas from the model quota and add fast rejection when the ceiling is reached.
- Add a PodDisruptionBudget, zone spread, workload identity and a default-deny egress policy.