A Kubernetes probe is a question the platform asks your pod on a timer, and it acts on the answer without asking a human. Get the question wrong and the platform will take confident, automated, fleet-wide action on bad information. For an agent service built with ADK Java, the classic mistakes are expensive. A liveness probe that calls the model restarts every replica during a provider brownout. A readiness probe that checks the session database removes every pod from the Service the moment the database blips. A missing startup probe kills slow-starting JVMs in a crash loop.

This page covers the semantics of the three probes, the timing arithmetic that decides how fast each one acts, how Spring Boot's availability states and health groups map onto them, how to write dependency checks that never block a probe, how readiness behaves during drain, and what changes on Cloud Run. The Kubernetes deployment article covers the warm-up runner and autoscaling. This one goes deeper on the probes themselves.

Three probes, three questions

Each probe asks one question and triggers one action:

ProbeQuestionAction on failureShould depend on
startupProbeHas the process finished starting?Kill and restart the container after the budgetBoot and warm-up only
livenessProbeIs this process wedged beyond self-repair?Kill and restart the containerInternal state of this JVM only
readinessProbeShould new requests be routed here right now?Remove the pod from Service endpoints; no restartLocal state and dependencies only this pod can fix

Liveness and readiness are disabled until the startup probe succeeds, which is what lets you give the startup probe a long budget without loosening liveness. The rule that prevents most outages follows directly from the action column. A probe should fail only for conditions that its action can fix. Restarting a pod does not fix an expired API key or a model quota. Removing one pod from the Service does not fix a database that every pod shares.

The timing arithmetic

Every probe has the same knobs, and Kubernetes defaults are periodSeconds 10, timeoutSeconds 1, failureThreshold 3, successThreshold 1 and initialDelaySeconds 0. successThreshold must be 1 for liveness and startup probes. An HTTP probe succeeds on any status from 200 to 399, and a request that exceeds timeoutSeconds counts as a failure.

The arithmetic that matters:

  • Time to act is roughly periodSeconds × failureThreshold, plus up to one timeout. With defaults, a dead pod is restarted or unrouted after 20 to 31 seconds.
  • Startup budget is failureThreshold × periodSeconds. A JVM that needs 90 seconds to load configuration, build agents and complete a warm-up run needs a budget comfortably above that, for example 36 × 5 = 180 seconds.
  • The 1-second default timeout is a trap for the JVM. A full GC pause, a saturated request thread pool or CPU throttling at the container limit can make a trivial endpoint take longer than a second. Three such slow answers in a row restart a healthy pod. Use 2 to 5 seconds, and serve probes from a path that does not queue behind agent traffic.

Spring Boot availability states and health groups

Spring Boot models this with two availability states. LivenessState is CORRECT or BROKEN. ReadinessState is ACCEPTING_TRAFFIC or REFUSING_TRAFFIC. With probes enabled, Actuator exposes /actuator/health/liveness and /actuator/health/readiness as health groups backed by those states. Spring Boot publishes ACCEPTING_TRAFFIC only after the application has started and every runner has completed, so a warm-up runner gates readiness automatically. Status DOWN and OUT_OF_SERVICE map to HTTP 503 and UP maps to 200.

# application.yaml
management:
  server:
    port: 8081                    # metrics and admin on a separate port
  endpoint:
    health:
      probes:
        enabled: true
        add-additional-paths: true  # also serve /livez and /readyz on the main port 8080
      group:
        liveness:
          include: livenessState
        readiness:
          include: readinessState
        dependencies:               # not used by any probe; scraped by monitoring
          include: modelProvider, sessionStore
          show-details: always
  endpoints:
    web:
      exposure:
        include: health, prometheus

The add-additional-paths setting, available since Spring Boot 2.6, matters when the management server runs on its own port. A probe against port 8081 tests the management server, which can stay healthy while the main server on 8080 is stuck. Probing /livez and /readyz on 8080 tests the connector that real traffic uses.

What goes in which probe for an agent

Which signal feeds which probe in an ADK Java agent servicestartupProbe/readyz, long budgetlivenessProbe/livez: is the JVM wedged?readinessProbe/readyz: send traffic here?warm-up runnerconfig, model name, one runlivenessStateBROKEN only on fatal statereadinessStatethis pod onlygatesbackground checkerevery 5 s, own timeoutcachedmodel providerNOT in any probedependencies groupdashboards and alertsProbes answer questions about this pod. Shared dependencies belong in alerts and circuit breakers.
Figure: probes read only local state or cached results. The model provider is monitored, never probed.

For an agent service the split is concrete.

  • Liveness includes only livenessState. Flip it to BROKEN from code when the process detects a state it cannot leave, for example an executor that has shut down or a corrupted in-memory registry. Never put a network call in it.
  • Readiness includes readinessState, which turns ACCEPTING_TRAFFIC once warm-up runners finish. Add a check only if it is truly local to this pod, such as a per-pod cache that must load first. A shared session database is not local: when it blips, every pod fails together.
  • The model provider is in neither. If Gemini or another provider is degraded, every replica sees the same failure. Failing readiness everywhere empties the Service, so even requests that never touch the model, such as session history or the health of your UI, start returning 503. Handle provider failures in the request path with timeouts, retries with backoff and a circuit breaker, and alert on them through the dependencies group and metrics.

Dependency checks that never block a probe

Shared dependencies such as the session database still need health indicators for the dependencies group and the aggregate /actuator/health endpoint. An indicator that performs I/O whenever it is called inherits the caller's timeout and multiplies load, since every pod checks every dependency for every scrape. Instead, check in the background on your own schedule and let the indicator return the last result, treating a stale result as a failure:

@Component
class SessionStoreHealthIndicator implements HealthIndicator {   // contributor name: "sessionStore"
    private static final Duration STALE_AFTER = Duration.ofSeconds(20);

    private final DataSource dataSource;
    private volatile Health last = Health.down().withDetail("reason", "not checked yet").build();
    private volatile Instant checkedAt = Instant.EPOCH;

    SessionStoreHealthIndicator(DataSource dataSource) { this.dataSource = dataSource; }

    @Scheduled(fixedDelay = 5000)            // needs @EnableScheduling on a configuration class
    void refresh() {
        long t0 = System.nanoTime();
        try (Connection c = dataSource.getConnection()) {
            boolean ok = c.isValid(2);        // driver-level check with its own 2 s timeout
            last = (ok ? Health.up() : Health.down())
                    .withDetail("latencyMs", (System.nanoTime() - t0) / 1_000_000).build();
        } catch (Exception e) {
            last = Health.down(e).build();
        }
        checkedAt = Instant.now();
    }

    @Override
    public Health health() {                  // called by the probe: no I/O, returns instantly
        if (Duration.between(checkedAt, Instant.now()).compareTo(STALE_AFTER) > 0) {
            return Health.down().withDetail("reason", "check is stale").build();
        }
        return last;
    }
}

Starting from DOWN rather than UNKNOWN matters, because UNKNOWN maps to HTTP 200 and would mark the pod ready before the first check. Treating staleness as DOWN catches the case where the scheduler thread itself has died. A ModelProviderHealthIndicator can follow the same pattern by reading your circuit breaker's state rather than calling the model.

To flip states from code, publish an event:

@Component
class AgentExecutorWatchdog {
    private final ApplicationEventPublisher events;
    AgentExecutorWatchdog(ApplicationEventPublisher events) { this.events = events; }

    void onExecutorTerminated(Throwable cause) {
        // Unrecoverable: let Kubernetes restart this container.
        AvailabilityChangeEvent.publish(events, this, LivenessState.BROKEN);
    }
}

Startup: slow JVMs and warm-up

An ADK Java service does real work before it can serve: it loads configuration, constructs agents and tools, and ideally resolves the model and runs one warm-up turn so that a bad model name or missing credential fails the rollout instead of the first user. Validation of that configuration is covered in the startup configuration validation article.

Give startup its own probe instead of inflating initialDelaySeconds on liveness. The startup probe polls frequently, so a fast boot becomes ready quickly, and a slow boot is still allowed its full budget. If the warm-up throws, the application exits, the new pod never becomes ready, and a rollout with maxUnavailable: 0 stalls while old pods keep serving. That is exactly the behavior you want from a bad deploy.

containers:
- name: agent
  ports: [{containerPort: 8080}]
  startupProbe:
    httpGet: {path: /readyz, port: 8080}   # green only after warm-up runners finish
    periodSeconds: 5
    failureThreshold: 36        # 180 s budget for boot plus warm-up
  livenessProbe:
    httpGet: {path: /livez, port: 8080}
    periodSeconds: 10
    timeoutSeconds: 3           # survive a GC pause
    failureThreshold: 3
  readinessProbe:
    httpGet: {path: /readyz, port: 8080}
    periodSeconds: 5
    timeoutSeconds: 2
    failureThreshold: 2
  lifecycle:
    preStop:
      exec: {command: ["sh", "-c", "sleep 10"]}   # needs a shell in the image
terminationGracePeriodSeconds: 120

Readiness during drain

When a pod is deleted, two things start in parallel: the kubelet runs the preStop hook and then sends SIGTERM, and the control plane removes the pod from EndpointSlices so load balancers and kube-proxy stop routing to it. Those updates propagate with a delay of a few seconds, so a pod that stops accepting connections the instant SIGTERM arrives still receives a short tail of new requests and fails them.

The preStop sleep covers that gap: the pod keeps serving normally while routing converges. Then SIGTERM triggers Spring's graceful shutdown, which stops accepting new requests and waits up to spring.lifecycle.timeout-per-shutdown-phase for in-flight ones. Agent turns that stream for a minute need that phase timeout, plus the preStop sleep, to fit inside terminationGracePeriodSeconds, or the kubelet sends SIGKILL mid-turn. Spring Boot also moves readiness to REFUSING_TRAFFIC when the context begins closing, but by then the endpoint removal is already under way, so do not rely on the readiness probe for drain. The details of draining streams are in the graceful shutdown article.

Cloud Run

Cloud Run supports startup, liveness and readiness probes on services. Startup probes gate the other two, and a failed readiness probe stops new traffic to that instance. Cloud Run's own documentation warns that it can route requests to a new instance before the first readiness check completes, so passing the startup probe must already mean the instance can serve. In practice that means doing the warm-up before the startup endpoint goes green, exactly as on Kubernetes. Cloud Run scales instances itself and manages its own request routing, so the preStop and endpoint-propagation mechanics above do not carry over. Deployment specifics are in the Cloud Run deployment article.

Worked example: a model provider outage

Six replicas serve an agent. At t = 0 the model provider starts returning errors and timeouts for most calls.

With the model in readiness (2-second timeout, period 5, threshold 2): by about t = 12 s every replica has failed twice. The Service has no ready endpoints, and the ingress returns 503 for everything, including the session list, transcript downloads and the status page that would have told users what was wrong. When the provider recovers, each pod needs another successful probe before traffic returns, and the six all come back within seconds of each other into a burst of retried traffic.

With the model in liveness: the same timeline, but each failure restarts the JVM. Pods spend the outage booting, failing warm-up because the model is down, and entering CrashLoopBackOff, whose back-off delay then keeps them down for minutes after the provider recovers.

With the design on this page: probes stay green. Agent turns fail fast once the circuit breaker opens and return a clear "model temporarily unavailable" message. The dependencies group reports modelProvider DOWN, and an alert fires on the error rate. Non-model endpoints keep working. When the provider recovers, the breaker's half-open probes restore service with no restarts.

Failure modes

  • Shared dependency in a probe: one external failure becomes a total outage or a restart storm.
  • Probe on the management port only: the main connector deadlocks while probes stay green.
  • Readiness used for load shedding: flipping REFUSING_TRAFFIC when in-flight runs pass a limit makes pods flap, and under a real spike all of them go unready at once. Return 429 from the application instead.
  • Default 1-second timeout: restarts under GC pressure, which adds load and worsens the pressure.
  • Indicator that does I/O per call: probe latency tracks dependency latency, and probe traffic becomes a measurable load on the database.
  • No startup probe: liveness kills slow boots, and the deployment never converges.

What to do next

  1. Write down, for each probe, the question it asks and the action it triggers. Remove anything from the probe that its action cannot fix.
  2. Enable add-additional-paths and point probes at /livez and /readyz on the main port.
  3. Convert every I/O health indicator to the cached, staleness-checked pattern, initialised to DOWN.
  4. Add a startup probe with a measured budget, and raise liveness and readiness timeouts above your worst observed GC pause.
  5. Check the shutdown arithmetic: preStop sleep plus shutdown phase timeout plus margin must be less than terminationGracePeriodSeconds.
  6. Game-day it: block egress to the model provider in staging and confirm that no pod restarts and no pod goes unready. Then wire the dependencies group into dashboards, as in the observability guide.
Key takeaway: A probe should fail only for conditions its action can fix. Liveness asks whether this JVM is wedged and reads only internal state. Readiness asks whether this pod should get traffic and reads local state plus cached per-pod checks. The startup probe gives boot and warm-up their own budget. Keep the model provider out of every probe, handle it with circuit breakers and alerts, and size timeouts for GC pauses.