An alert is a request for a person's attention, and an agent fleet can generate requests faster than any team can read them. A model provider wobbles for ten minutes and forty agents each report three symptoms; a prompt change doubles tool calls and nothing pages because no single call failed. Good alerting for agents is mostly about routing, identity and restraint, not about inventing more signals.
This article covers the delivery side. Which SLIs to pick and how burn-rate thresholds work are in Defining SLOs for LLM Agents; what to do once paged is in the incident response playbook. Here you build an ADK Java plugin that raises event alerts, push them to Prometheus Alertmanager correctly, and configure grouping and inhibition so one root cause produces one page. Plugin signatures come from the google-adk 1.11.0 jar; Alertmanager fields from its v2 OpenAPI spec and configuration reference.
Rate alerts and event alerts
Agent alerts come in two shapes, and mixing them up causes most alerting pain.
| Rate alerts | Event alerts | |
|---|---|---|
| Question | is the fleet worse than it should be? | did one specific bad thing happen? |
| Examples | error ratio, p95 latency, tokens per run, SLO burn | a write tool ran twice in one invocation; a guardrail blocked an exfiltration attempt; a run hit its LLM-call cap |
| Source | metrics, evaluated by Prometheus rules | callbacks, pushed directly |
| Needs a threshold over time | yes, with for: or burn windows | no, one occurrence is enough |
| Resolves when | the expression stops matching | a hold time passes |
Default to rate alerts. Emit counters from callbacks, as in tool observability metrics, and let the rule engine decide; it sees the whole fleet, survives agent restarts and can be tested offline. Push an event alert directly only when a single occurrence is actionable and a counter would bury it: the duplicate refund matters even if it is one in a million calls.
Architecture
Both sources meet at Alertmanager, which is the only component that decides who is notified. Agents never call a pager directly; that would bypass grouping, silences and inhibition, the three tools that keep a storm to one page.
Detecting events in a plugin
A plugin sees every agent, model and tool callback for every run on the runner, which makes it the right place for event detection. Two rules: never block the agent, and never throw into it. Detection is a few map lookups; delivery happens on another thread.
public final class AlertingPlugin extends BasePlugin {
private final AlertSink sink;
private final Set<String> writeTools; // e.g. "issue_refund"
private final String region; // from deployment config
private final Map<String, Map<String, Integer>> writes = new ConcurrentHashMap<>();
public AlertingPlugin(AlertSink sink, Set<String> writeTools, String region) {
super("alerting");
this.sink = sink;
this.writeTools = writeTools;
this.region = region;
}
@Override
public Maybe<Map<String, Object>> beforeToolCallback(
BaseTool tool, Map<String, Object> args, ToolContext ctx) {
if (writeTools.contains(tool.name())) {
int n = writes.computeIfAbsent(ctx.invocationId(), k -> new ConcurrentHashMap<>())
.merge(tool.name(), 1, Integer::sum);
if (n == 2) {
sink.offer(Alert.event("AgentWriteToolRepeated", "page",
Map.of("region", region, "agent", ctx.agentName(), "tool", tool.name()),
Map.of("summary", tool.name() + " called twice in one invocation",
"invocation", ctx.invocationId(), "session", ctx.sessionId(),
"runbook_url", "https://runbooks.example.com/agent-write-repeated")));
}
}
return Maybe.empty(); // never alter the call here
}
@Override
public Completable onRunErrorCallback(InvocationContext ctx, Throwable t) {
if (t instanceof LlmCallsLimitExceededException) {
sink.offer(Alert.event("AgentLlmCallCapHit", "ticket",
Map.of("region", region, "agent", ctx.agent().name()),
Map.of("summary", "run exceeded maxLlmCalls", "invocation", ctx.invocationId())));
}
writes.remove(ctx.invocationId()); // failed runs clean up too
return Completable.complete();
}
@Override
public Completable afterRunCallback(InvocationContext ctx) {
writes.remove(ctx.invocationId()); // no per-invocation leak
return Completable.complete();
}
}Register it with new InMemoryRunner(agent, APP, List.of(alerting)) or the plugin list of the full Runner constructor. LlmCallsLimitExceededException is what InvocationContext throws when RunConfig.maxLlmCalls is exceeded; confirm in your version that it reaches onRunErrorCallback rather than ending as an error event, and handle both if in doubt. Detection does not replace blocking: if a repeated write must never happen, a guardrail callback should refuse it as well. The alert tells people it was attempted.
The AlertSink is a bounded queue drained by one thread. When it is full it drops the alert and increments alerts_dropped_total; it never waits. An agent that stalls because the alerting system is slow has turned a monitoring problem into an outage.
The push contract
Alertmanager accepts alerts on POST /api/v2/alerts as a JSON array. Each alert has a required labels map and optional annotations, startsAt, endsAt and generatorURL. Three rules follow from how Alertmanager treats them:
- Labels are identity. Alerts with the same label set are the same alert; posting it again updates it rather than creating a new one. Grouping and routing also work on labels. So labels hold low-cardinality facts: alert name, severity, agent, tool, region. Session and invocation ids go in annotations. Put a session id in a label and every event becomes a separate alert that groups with nothing.
- Every alert needs an end. An alert without
endsAtis resolved by Alertmanager afterresolve_timeout(5 minutes by default) unless it is posted again. Rule engines keep re-sending while a condition holds; an event has no condition to hold, so setendsAtexplicitly to a hold time, such as 30 minutes after the event. A repeat of the event inside that window re-posts and extends it. - Send to every replica. In a high-availability pair, Alertmanager instances share state and de-duplicate notifications among themselves. Clients should post to each instance, not through a load balancer that picks one.
// drained by one thread every 2 s; batches whatever is queued
void flush(List<Alert> batch) {
String body = json.writeValueAsString(batch.stream().map(a -> Map.of(
"labels", a.labels(), // alertname, severity, agent, tool
"annotations", a.annotations(), // summary, runbook_url, session, invocation
"startsAt", a.at().toString(),
"endsAt", a.at().plus(a.hold()).toString(),
"generatorURL", "https://traces.example.com/inv/" + a.annotations().get("invocation")))
.toList());
for (URI am : alertmanagers) { // every replica, not a load balancer
HttpRequest req = HttpRequest.newBuilder(am.resolve("/api/v2/alerts"))
.timeout(Duration.ofSeconds(3))
.header("Content-Type", "application/json")
.POST(HttpRequest.BodyPublishers.ofString(body)).build();
http.sendAsync(req, HttpResponse.BodyHandlers.discarding())
.exceptionally(e -> { failed.increment(); return null; });
}
}
Grouping, routing and inhibition
Routing turns alerts into notifications. Group by the labels that describe one problem, page only for severities that need a human now, and use inhibition so a known root cause silences its symptoms.
route:
receiver: agent-tickets
group_by: [alertname, agent]
group_wait: 30s # defaults shown explicitly: first notification delay
group_interval: 5m # wait before notifying about changes to a group
repeat_interval: 4h # re-notify an unchanged, still-firing group
routes:
- matchers: [severity="page"]
receiver: agent-oncall
- matchers: [alertname="Watchdog"]
receiver: deadman # an external service that pages if this STOPS arriving
repeat_interval: 1m
inhibit_rules:
- source_matchers: [alertname="ModelProviderDegraded"]
target_matchers: [alertname=~"AgentToolErrorRatioHigh|AgentLatencyHigh"]
equal: [region] # both sides must carry the same region labelThe provider alert comes from a rate rule on model error classes, for example 5xx or 429 responses over all agents, aggregated by region. Inhibition compares the equal labels, so every rule and every pushed alert must carry region; an aggregation that drops it silently disables the rule. While the provider alert fires, per-agent error and latency alerts in that region are inhibited: one page that says the provider is degraded is more useful than forty that say each agent is slow. Inhibition has a cost: the equal list must be right, or it hides independent failures. Never inhibit event alerts such as AgentWriteToolRepeated or AgentLlmCallCapHit; a provider outage does not explain a duplicate refund. Two repeats with the same labels inside the hold window merge into one alert and the later annotations win, so point generatorURL at a query over all matching events rather than at one trace.
The Watchdog alert is a rule whose expression is always true. It proves the whole path works. If the external dead-man service stops receiving it, Prometheus, Alertmanager or the network between them is broken, and you learn that before you need an alert that cannot arrive.
Rules with tests
Metric alerts are configuration, and configuration deserves tests. A tool error-ratio rule with a for: clause, a runbook link and a unit test:
# rules/agent.yml
groups:
- name: agent
rules:
- alert: AgentToolErrorRatioHigh
expr: |
sum by (region, agent, tool) (rate(agent_tool_calls_total{outcome="error"}[10m]))
/ sum by (region, agent, tool) (rate(agent_tool_calls_total[10m])) > 0.05
for: 10m
labels: {severity: ticket}
annotations:
summary: "{{ $labels.tool }} failing for {{ $labels.agent }}"
runbook_url: https://runbooks.example.com/agent-tool-errors
# tests/agent_test.yml, run with: promtool test rules tests/agent_test.yml
rule_files: [../rules/agent.yml]
evaluation_interval: 1m
tests:
- interval: 1m
input_series:
- series: 'agent_tool_calls_total{region="eu",agent="billing",tool="lookup",outcome="error"}'
values: '0+6x30'
- series: 'agent_tool_calls_total{region="eu",agent="billing",tool="lookup",outcome="ok"}'
values: '0+60x30'
alert_rule_test:
- eval_time: 25m
alertname: AgentToolErrorRatioHigh
exp_alerts:
- exp_labels: {severity: ticket, region: eu, agent: billing, tool: lookup}
exp_annotations:
summary: "lookup failing for billing"
runbook_url: https://runbooks.example.com/agent-tool-errorsThe metric names are your own, emitted by your plugin, not ADK built-ins. Test the routing tree too: amtool config routes test severity=page agent=billing prints the receiver an alert with those labels would reach. Run both in CI, so a typo in a matcher fails a build instead of silencing a pager.
Worked example: a provider storm and one real page
A Tuesday afternoon, forty agents in one region. At 14:02 the model provider starts returning 503 for a fifth of requests. Without routing design, each agent's error-ratio alert, latency alert and LLM-call-cap ticket fire separately: 120 notifications in four minutes, and the on-call engineer spends the first quarter hour working out that they are one problem.
With the configuration above, ModelProviderDegraded fires at 14:07 after its for: 5m. The per-agent symptom rules use for: 10m, so none of them has fired yet; when they do, a matching provider alert in the same region is already active and they are inhibited from the start. The team receives one page. Keeping the provider rule's for: shorter than the symptoms' is what makes this work, a deliberate ordering worth writing down next to the rules. The LLM-call-cap events are not inhibited; they arrive as tickets grouped per agent, which is acceptable because tickets wake nobody.
At 14:31, while the provider alert is still active, a retry wrapper re-executes a timed-out issue_refund call. The plugin sees the second call in the same invocation and pushes AgentWriteToolRepeated with the session id in its annotations. It is not inhibited, it pages, and the engineer follows generatorURL to the traces. That is the alert that saved money; the forty agents being slow was something the provider was already fixing.
Failure modes
| Failure | Symptom | Fix |
|---|---|---|
| High-cardinality labels | every event its own alert, no grouping, huge Alertmanager memory | ids in annotations, labels from a fixed set |
| Sink blocks the agent | turn latency rises when Alertmanager is slow | bounded queue, drop and count, async posting |
| Events never resolve or resolve instantly | stale pages or alerts gone before anyone looks | explicit endsAt hold time |
| Silent alerting path | no alerts during an outage, and nobody notices | always-firing Watchdog to an external dead-man service |
| Over-broad inhibition | real failures hidden behind a provider alert | tight equal labels; never inhibit event alerts |
| Flapping | fire and resolve every few minutes | longer for:, hysteresis on thresholds |
| Pages nobody can act on | acknowledged and ignored | every page has a runbook and an owner, or becomes a ticket |
Hygiene and trade-offs
Alerting decays. Each month, list every alert that fired with how many times, whether someone acted, and how long it took. An alert that fired often and led to no action is either mis-thresholded or belongs on a dashboard. A page should mean a user is hurt now or soon; everything else is a ticket. Add an owner label so routing and reviews have someone to ask.
Trade-offs. Direct event pushes are fast and precise but bypass the rule engine's fleet-wide view and its tests, so keep their number small. Metric rules are robust and testable but cost a scrape and evaluation interval in detection time. Inhibition cuts noise and can hide independent failures. Shorter group_wait shortens time to first notification and splits a storm into more messages. The broader picture of what to instrument is in ADK Java observability.
What to do next
- List your current alerts and classify each as rate or event; move rate signals into Prometheus rules.
- Add the alerting plugin for the two or three events that matter individually, starting with repeated write tools.
- Make the sink bounded and asynchronous, count drops, and post to every Alertmanager replica.
- Audit labels: no session, invocation or user ids.
- Set
endsAton every pushed event. - Add a provider-degraded rule and an inhibition rule with a tight
equallist. - Deploy an always-firing Watchdog to an external dead-man service.
- Put
promtool test rulesandamtool config routes testin CI. - Book a monthly review of what fired and what anyone did about it.