A liveness probe tells you the JVM answers HTTP. A dashboard tells you how many turns completed and how long they took. Neither tells you that the agent answered a refund question politely, quickly, with status 200, and without ever calling the refund-status tool, because last night a dependency renamed a field and the tool now returns an error the model papers over. Agents fail semantically while every infrastructure signal stays green.

Synthetic monitoring closes that gap. You keep a small library of scripted conversations with known correct behaviour, run them against production on a schedule from outside the service, and assert not only that a reply arrived but that the agent took the right path to it. This article builds that system for an ADK Java agent: what a probe asserts, how to isolate synthetic users and their side effects, how to size cadence against cost, and how to alert without paging for noise. It assumes you already ship the standard metrics and dashboards and complements, rather than replaces, continuous evaluation of real traffic.

What probes catch that nothing else does

Each production signal answers a different question, and synthetic probes earn their cost only for the questions nothing else answers. The table places them.

SignalQuestion it answersBlind spot
Liveness and readinessIs the process up and able to take traffic?Says nothing about model, tools or answers
Deploy smoke testDid this revision complete one real turn before traffic?Runs once per deploy, not while the world changes
Metrics and dashboardsHow many turns, how slow, how many errors?A wrong answer with status 200 counts as success
Continuous evaluationHow good are sampled real conversations?Delayed, statistical, and only covers what users happened to ask
Synthetic probesDoes a known task still follow its known path, right now, in every region?Only covers the scripted tasks; costs tokens on every run

The distinctive property is the known answer. Because you wrote the conversation and control the fixture data behind it, you know which tools should be called, with which arguments, and which facts the final reply must contain. That turns a fuzzy quality question into a crisp pass or fail, available every few minutes, including at 3 a.m. when real traffic is too thin for any statistical signal. The deploy pipeline's smoke test uses the same idea once per revision; synthetic monitoring repeats it continuously because most agent breakages are not caused by your deploys: upstream APIs change, model aliases shift, quotas are cut.

Anatomy of a probe

A synthetic probe, from schedule to alertProbe schedulerper region, jitteredPublic agent APIsame path as usersADK Java RunnerLlmAgent + toolsReal modelreal quotaturnSynthetic guardbeforeToolCallbackSandbox tenantwrite tools land hereuser synthetic-*Transcriptcalls, errors, texteventsAssertions L0 transport, L1 protocol,L2 trajectory, L3 contentAlert rulespage on L0/L1, ticket on L2/L3 ratesThe probe travels the user path end to end; only the side effects are diverted.
Probes use the user path end to end; a guard diverts write tools to a sandbox for synthetic identities.

A probe is data, not code. Keeping probes declarative means the people who own an agent's behaviour can add one without touching the runner, and every run produces a comparable record. A useful probe definition carries the turns to send, the fixture identity, and assertions in four layers that fail for different reasons and deserve different responses.

{
  "id": "refund-status-happy-path",
  "owner": "payments-agent",
  "user": "synthetic-refund-01",
  "turns": ["What is the status of my refund for order SYN-4471?"],
  "assert": {
    "transport":  { "status": 200, "maxLatencyMs": 20000 },
    "protocol":   { "noErrorCode": true, "finalResponse": true },
    "trajectory": { "mustCall": [ { "tool": "get_refund_status",
                                    "args": { "orderId": "SYN-4471" } } ],
                    "mustNotCall": ["issue_refund"], "maxToolCalls": 4 },
    "content":    { "mustContain": ["SYN-4471", "processed"],
                    "mustNotMatch": ["(?i)couldn.t find", "(?i)error"] }
  }
}
  • L0 transport: status code and latency. A failure here is infrastructure: network, load balancer, overload, cold start.
  • L1 protocol: the event stream ended with a final response and no event carried an errorCode(). A failure here usually means a model-side block, a quota error or an exception inside the run.
  • L2 trajectory: the expected tools were called with the expected arguments, forbidden tools were not, and the number of calls stayed inside a budget. A failure here is behavioural: prompt, tool schema or model drift.
  • L3 content: deterministic checks on the reply text, such as required identifiers and forbidden phrases. Keep these narrow; wording changes legitimately.

Notice what L3 does not do: it does not call an LLM judge. A judge on every probe run doubles cost and adds its own nondeterminism to the one signal you want crisp. Put judged quality in continuous evaluation and keep probes to checks a regular expression or a set comparison can decide.

Isolating synthetic traffic

Probes travel the real user path, so without care they create real side effects and real data. Isolation has three parts. First, identity: every probe uses a user id with a reserved prefix such as synthetic-, created only for probes and tied to fixture records (order SYN-4471 exists in the payments sandbox and nowhere else). Second, data hygiene: analytics, billing, evaluation sampling and any memory ingestion filter that prefix out, so probes never become someone's training example or a usage spike. Third, side effects: tools that write to the outside world must not fire for synthetic users.

The cleanest enforcement point in ADK Java is a before-tool callback. In 1.11.0, Callbacks.BeforeToolCallbackSync receives the InvocationContext, the BaseTool, the argument map and the ToolContext, and returns Optional<Map<String, Object>>: an empty optional lets the tool run, a present map is used as the tool's result instead. That lets one guard divert every write tool for synthetic users without touching tool code.

import com.google.adk.agents.Callbacks;
import java.util.Map;
import java.util.Optional;
import java.util.Set;

public final class SyntheticGuard {
  private static final Set<String> WRITE_TOOLS = Set.of("issue_refund", "send_email", "cancel_order");

  public static Callbacks.BeforeToolCallbackSync create(SandboxClient sandbox) {
    return (invocation, tool, args, toolContext) -> {
      if (!toolContext.userId().startsWith("synthetic-")) {
        return Optional.empty();                       // real user: run the real tool
      }
      if (WRITE_TOOLS.contains(tool.name())) {
        // Exercise the write path against the sandbox tenant and return its result.
        return Optional.of(sandbox.execute(tool.name(), args));
      }
      return Optional.empty();                         // reads run for real against fixtures
    };
  }
}

// Wiring: LlmAgent.builder()...beforeToolCallbackSync(SyntheticGuard.create(sandbox)).build();

The guard routes writes to a sandbox tenant rather than returning a canned success: a canned map proves only that the model chose the right tool and arguments, a sandbox call also proves the write path works. Reads hit real systems against fixture records, so a renamed field in the real refund API breaks the probe. Add a second check inside each write tool that refuses fixture identifiers such as order ids starting SYN-, in case a refactor ever bypasses the prefix check.

The probe runner

The runner is ordinary Java: a scheduler, an HTTP client that calls your public agent API exactly as a client application would, and an evaluator. The API shape is yours, so the sketch below assumes your service returns the turn's events in some form you can reduce to a transcript of tool calls, error codes and final text.

record ToolCall(String name, Map<String, Object> args) {}
record Transcript(int status, long latencyMs, List<ToolCall> calls,
                  List<String> errorCodes, boolean hasFinal, String finalText) {}

final class ProbeRunner {
  private final ScheduledExecutorService pool = Executors.newScheduledThreadPool(4);

  void start(List<Probe> probes, Duration every) {
    for (Probe probe : probes) {
      long jitterMs = ThreadLocalRandom.current().nextLong(every.toMillis());
      pool.scheduleAtFixedRate(() -> runOnce(probe), jitterMs, every.toMillis(), MILLISECONDS);
    }
  }

  private void runOnce(Probe probe) {
    String sessionId = "probe-" + UUID.randomUUID();      // fresh session every run
    Transcript t;
    try {
      t = client.runTurns(probe.user(), sessionId, probe.turns());
    } catch (Exception e) {
      results.record(probe, Layer.TRANSPORT, false, e.toString());
      return;
    }
    results.record(probe, Layer.TRANSPORT, checkTransport(probe, t), "");
    results.record(probe, Layer.PROTOCOL,  t.errorCodes().isEmpty() && t.hasFinal(), "");
    results.record(probe, Layer.TRAJECTORY, checkTrajectory(probe, t), "");
    results.record(probe, Layer.CONTENT,   checkContent(probe, t), "");
  }

  static boolean checkTrajectory(Probe p, Transcript t) {
    boolean required = p.mustCall().stream().allMatch(want -> t.calls().stream().anyMatch(got ->
        got.name().equals(want.tool()) && got.args().entrySet().containsAll(want.args().entrySet())));
    boolean forbidden = t.calls().stream().anyMatch(got -> p.mustNotCall().contains(got.name()));
    return required && !forbidden && t.calls().size() <= p.maxToolCalls();
  }
}

A fresh session per run stops earlier conversations steering the answer, jitter spreads load, and subset argument matching lets the model add optional arguments. For an in-process variant in staging, the same transcript can be built straight from the Flowable<Event> that Runner.runAsync(userId, sessionId, content) returns: collect event.functionCalls(), any present event.errorCode(), and the text of the event where event.finalResponse() is true. In production stay black-box; an in-process probe skips the network, authentication and load balancer.

Cadence, detection time and cost

Cadence trades detection time against cost and against load you add to your own dependencies. Time to detect is roughly the probe interval multiplied by the number of consecutive failures your alert requires, plus one run's latency. Cost is linear in probes, regions and frequency, and every run spends real model tokens.

InputValueResult
Probes6
Regions probed from3
Interval5 minutes288 runs per probe per region per day
Runs per day6 x 3 x 2885,184
Model calls per runabout 2 (tool call, then answer)10,368 calls per day
Input tokens per callabout 3,000 (instruction, tools, history)about 31 million input tokens per day
Detection time, 2 consecutive failures2 x 5 min + run latencyabout 10 to 11 minutes

Multiply the token line by your model's price per million tokens before committing; the answer usually argues for tiers: two or three critical-path probes every few minutes, the broader library every 30 to 60 minutes. Count probe traffic against your quotas, and give the runner a kill switch, because during an incident a probe fleet hammering a struggling dependency is part of the problem.

Alerting on a non-deterministic system

Probes against an LLM agent are not perfectly repeatable: the same prompt can take a different valid path, and single runs fail on transient model errors. Treat the layers differently, because their failures have different base rates and urgency.

LayerAlert conditionAction
L0 transportSame probe fails 2 consecutive runs in 2 or more regionsPage
L0 transportFails in exactly 1 regionPage the regional owner only
L1 protocol3 or more failures across probes in 15 minutesPage; check quota and model status
L2 trajectoryFailure rate above 20 percent for one probe over 1 hourTicket to agent owner, high priority
L3 contentFailure rate above 30 percent over 3 hoursTicket; review the probe and the agent

Measure each probe's baseline flake rate during a quiet week before you set thresholds, and record it next to the probe. A probe failing 8 percent of the time on a healthy system has assertions that are too tight; fix the probe, not the threshold. Attach the failing transcript to every alert so on-call sees which tool was or was not called. The incident response playbook picks up from there.

Worked example: the fluent apology

A payments agent runs six probes from three regions every five minutes. At 02:10 the upstream refunds API ships a change that renames status to refundState in its response. The refund-status tool now returns {"error": "missing status"}.

  1. 02:12, first run after the change: L0 passes (200 in 3.1 s), L1 passes (no error code, final response present), L2 passes (get_refund_status called with SYN-4471). L3 fails: the reply says "I could not retrieve the status of your refund right now."
  2. 02:17 and 02:22: the same result in all three regions. The L3 rate for this probe crosses 30 percent within the window, but the rule needs three hours of data, so nothing pages yet.
  3. 02:25: the team had added a fifth assertion after a previous incident, a tool-response check that no get_refund_status response contains an error key. It is classed as L2 and its threshold is reached after three consecutive runs. A high-priority ticket opens with the transcript attached.
  4. The on-call engineer reads the transcript, sees the tool error, finds the upstream change, and ships a tool fix that accepts both field names.

The lesson the team wrote down: assert on tool responses, not only on tool calls. Agents handle tool errors gracefully, so a failing dependency produces a fluent apology rather than an exception; without the probe, the first signal would have been the next morning's support queue.

Failure modes

  • Fixture rot. Someone cleans up "test data" and SYN-4471 disappears; every probe fails at once. Own fixtures in code, recreate them on a schedule, and alert on fixture checks separately from agent checks.
  • Probes that pin wording. A model upgrade rephrases replies and L3 fails everywhere. Assert on identifiers and facts, never on sentences.
  • Pollution. Probe sessions leak into evaluation samples, memory or business metrics. Filter the prefix at every consumer.
  • Global-only probing. One vantage point hides regional outages and quotas.
  • Silent probe death. A crashed runner looks like all-green. Alert on missing results.
  • Self-inflicted load. Probes keep firing at full rate during an incident. Back off automatically after repeated L0 failures.

Trade-offs

ChoiceOption AOption B
Write toolsCanned result: cheap, proves intent onlySandbox tenant: proves the write path, needs upkeep
AssertionsDeterministic: crisp, may miss nuanceLLM judge: nuanced, costly and noisy
CadenceMinutes: fast detection, real token costHourly: cheap, slow detection
VantageBlack-box via public API: covers everything users hitIn-process Runner: faster, skips network and auth
Library sizeFew critical paths: low noiseMany tasks: wider coverage, more maintenance and flakes

What to do next

  1. List the three tasks whose silent failure would hurt most, and write one probe definition per task with all four assertion layers.
  2. Create synthetic- users and fixture records in a sandbox tenant, owned in code and recreated on a schedule.
  3. Add a before-tool callback that diverts write tools for synthetic users, plus a fixture-id refusal inside each write tool.
  4. Filter the synthetic prefix out of analytics, billing, evaluation sampling and memory ingestion, and add a test that proves the filter works.
  5. Run the probes for a week without alerting to measure flake rates, then set per-layer thresholds and a missing-results alert.
  6. Compute daily token cost for your cadence and split probes into a fast and a slow tier.
  7. Attach the transcript to every alert and link the runbook from it.
  8. Review your liveness and readiness probes so process health and semantic health are alerting separately.
Key takeaway: Infrastructure signals stay green while agents fail semantically, so run scripted conversations with known answers against production on a schedule. Assert in layers: transport, protocol, trajectory and content, and include tool responses, because agents apologise fluently over broken tools. Give probes reserved identities, divert their writes with a before-tool callback, filter them from every downstream consumer, size cadence against token cost, and alert per layer with thresholds measured on a healthy week.