Server-side metrics tell you what the inference engine believes happened to the requests it received. They say nothing about requests that never arrived because DNS, TLS, a gateway or a rate limiter failed first, and nothing about a model that answers quickly with garbage. Synthetic monitoring closes that gap: a small program sends known prompts to your LLM endpoint on a schedule, from the places your users are, and records how long the first token took, how the tokens streamed and whether the answer passed a check you wrote in advance.
This article builds such a prober for a GPU-backed LLM service. It covers what to measure on the client side, how to design probe prompts so caching and batching do not fool you, a working Python prober that exports Prometheus metrics, how to target individual replicas, what probes cost on the GPUs, and how to alert on them without paging on every hiccup. The server-side view is covered in LLM SLO burn rate alerts; this page is the outside-in half.
What only an outside probe can see
There are three classes of failure that only an outside probe catches reliably. The first is the path in front of the engine: an expired certificate, a misrouted DNS record, an API gateway that rejects a new auth token format, or a load balancer that sends traffic to a pod that has not finished loading weights. The engine sees no requests, so its error rate is zero and its latency histograms look excellent.
The second is silence. A low-traffic model at three in the morning has no requests, so a broken deployment produces no bad events and no alert until the first real user arrives. A probe produces traffic on a schedule, which turns an unobservable state into a measured one.
The third is wrong answers delivered on time. A bad tokenizer file, a quantised checkpoint with a broken scale, a chat template that drops the system prompt, or a kernel regression after an engine upgrade can all return HTTP 200 with fluent nonsense. Latency and error metrics are blind to this. A probe with a verifiable expected answer is the cheapest detector you can run continuously.
What to measure from the client
A streaming probe should record the same quantities users feel, measured from the client clock. Start the timer immediately before sending the request, so connection setup and TLS are inside the measurement, then record:
- Time to first token (TTFT): until the first content-bearing chunk arrives. Role-only or empty chunks do not count; some servers send one before any text.
- Inter-chunk gaps: the time between successive content chunks. One server-sent event can carry more than one token, so this is not the same as inter-token latency; divide the decode time by the token count for a mean time per output token instead.
- Total time and output tokens: tokens come from the usage block if the server returns one. The OpenAI API defines
stream_options: {"include_usage": true}for streamed usage; check that your serving stack and version honour it before you depend on it, and fall back to counting with the model tokenizer. - Outcome: HTTP status, whether the stream ended with a finish reason, and whether the content passed the probe check.
The gap between probe TTFT and the engine's own TTFT histogram is itself a signal. If the engine reports 180 ms and the probe from the same region sees 900 ms, the time is going to the gateway, the network or queueing in front of the engine, and no amount of GPU tuning will fix it. The arithmetic of the engine-side number is in the TTFT ledger.
Where probes run
Run probes from at least two vantage points outside your network and one inside the cluster. External runners exercise the full path. The in-cluster runner sends the same probes straight to each replica's pod address, bypassing the load balancer, so a single bad GPU node shows up as one failing series instead of a fraction of public probes failing at random. Most fleets discover replicas through the orchestrator API and refresh the target list every minute.
Designing probe prompts
Probe prompts are test cases, so design them deliberately. Four probe types cover most fleets:
| Probe | Prompt shape | What it catches |
|---|---|---|
| liveness | 20 tokens in, 8 out | path failures, cold pods, auth |
| long prefill | about 4,000 tokens in, 16 out | prefill regressions, context limits, KV pressure |
| long decode | 50 in, 256 out | decode speed, stream stalls, preemption |
| correctness | task with a checkable answer | bad weights, templates, tokenizers |
Three details decide whether the numbers mean anything. Defeat prefix caching. Engines such as vLLM reuse KV blocks for repeated prefixes, so a fixed 4,000-token probe becomes a cache hit after the first run and measures almost no prefill. Put a random nonce at the very start of the prompt, before any shared text, when you want to measure cold prefill; keep a second, cached variant if you want to watch cache hit behaviour, as described in prefix caching.
Fix the output length. Set max_tokens and use a prompt whose natural answer is longer, so the model does not stop early and shorten the decode measurement. vLLM accepts an extra ignore_eos sampling parameter for this; it is not part of the OpenAI API, so only use it where you have confirmed support.
Check properties, not strings. Greedy decoding with temperature 0 is not guaranteed to be bitwise reproducible on GPUs, because batch composition changes the order of floating point reductions. Check that a JSON answer parses and matches a schema, that an arithmetic answer equals a computed value, or that a required keyword appears. Exact string comparison will flap.
A working prober
The prober below sends a streamed chat completion with httpx, timestamps each chunk, runs the check, and exports Prometheus metrics on port 9105. It assumes an OpenAI-compatible endpoint.
import json, os, random, re, string, time
import httpx
from prometheus_client import Counter, Gauge, Histogram, start_http_server
LAT = Histogram("probe_ttft_seconds", "client TTFT", ["probe", "target"],
buckets=(.1, .2, .4, .8, 1.6, 3.2, 6.4, 12.8))
TPOT = Gauge("probe_time_per_chunk_seconds", "decode s/chunk", ["probe", "target"])
OUTCOMES = ("ok", "wrong_answer", "incomplete", "transport_error")
RUNS = Counter("probe_runs_total", "probe outcomes", ["probe", "target", "outcome"])
def number_after(key, text):
m = re.search(r'"%s"\s*:\s*(-?\d+)' % key, text) # tolerates code fences
return int(m.group(1)) if m else None
PROBES = {
"liveness": {
"messages": [{"role": "user", "content": "Reply with the word ready."}],
"max_tokens": 8,
"check": lambda text: bool(text.strip()),
},
"correctness": {
"messages": [{"role": "user", "content":
"Reply with only a JSON object {\"sum\": N} where N is 1847 + 2765."}],
"max_tokens": 32,
"check": lambda text: number_after("sum", text) == 4612,
},
}
def nonce():
return "".join(random.choices(string.ascii_lowercase, k=12))
def run_probe(base_url, target, name, spec, model, timeout=30.0):
msgs = [dict(m) for m in spec["messages"]]
msgs[0]["content"] = f"[probe {nonce()}] " + msgs[0]["content"]
body = {"model": model, "messages": msgs, "max_tokens": spec["max_tokens"],
"temperature": 0, "stream": True}
headers = {"Authorization": f"Bearer {os.environ['PROBE_TOKEN']}",
"X-Synthetic-Probe": name}
t0 = time.perf_counter()
first, chunks, text, finished = None, 0, [], False
try:
with httpx.Client(timeout=timeout) as client:
with client.stream("POST", f"{base_url}/v1/chat/completions",
json=body, headers=headers) as r:
if r.status_code != 200:
RUNS.labels(name, target, f"http_{r.status_code}").inc()
return
for line in r.iter_lines():
if not line.startswith("data: ") or line == "data: [DONE]":
continue
ev = json.loads(line[6:])
for ch in ev.get("choices", []):
piece = ch.get("delta", {}).get("content")
if piece:
if first is None:
first = time.perf_counter()
chunks += 1
text.append(piece)
if ch.get("finish_reason"):
finished = True
except (httpx.HTTPError, json.JSONDecodeError):
RUNS.labels(name, target, "transport_error").inc()
return
end = time.perf_counter()
if first is None or not finished:
RUNS.labels(name, target, "incomplete").inc()
return
LAT.labels(name, target).observe(first - t0)
if chunks > 1:
TPOT.labels(name, target).set((end - first) / (chunks - 1))
try:
ok = spec["check"]("".join(text))
except Exception:
ok = False
RUNS.labels(name, target, "ok" if ok else "wrong_answer").inc()
if __name__ == "__main__":
for name in PROBES:
for o in OUTCOMES:
RUNS.labels(name, "public", o) # pre-create so increase() sees the first failure
start_http_server(9105)
while True:
for name, spec in PROBES.items():
run_probe(os.environ["PROBE_URL"], "public", name, spec, os.environ["PROBE_MODEL"])
time.sleep(30 + random.uniform(-3, 3))Note the decode figure is per chunk, not per token, for the reason given above; switch to the usage block when your server provides it. The jittered sleep stops a fleet of probers from hitting the endpoint in lockstep. A wrong answer is recorded as its own outcome, never folded into errors, because it needs a different responder.
Worked example: what probes cost
Probes run on the same GPUs as users, so size them. Take three locations, four probe types and one run every 30 seconds: 3 x 4 x 2 = 24 probes per minute, or 34,560 per day. With an average of about 1,000 prompt tokens (the long-prefill probe dominates) and 80 output tokens, that is about 35 million prefill tokens and 2.8 million decode tokens per day.
On a fleet serving billions of tokens a day this is noise. On a single 8-GPU node serving a niche model it can be a visible slice of decode capacity, and every probe occupies a batch slot. Two rules keep it honest: lower frequency for expensive probes (the long-decode probe every five minutes is plenty), and tag every probe with a header such as X-Synthetic-Probe so the gateway can exclude probe traffic from user-facing SLO counters and from billing. Without the tag, probes inflate request counts on quiet models and their short prompts drag latency percentiles down, hiding real regressions.
Alerting on probe results
Single probe failures are common and mostly meaningless: a packet loss, a pod that was draining. Alert on patterns instead. A reasonable starting set:
- Path down: the liveness probe failed at least three of the last four runs in at least two of three external locations. Quorum across locations separates your outage from a runner's network problem.
- Replica bad: one replica's in-cluster probe failed or ran at more than three times the fleet median TTFT for five minutes. Route this to automation that cordons the pod, not to a human.
- Wrong answers: any correctness failure on two consecutive runs pages immediately. This is the rarest alert and the most valuable; it usually means a bad rollout, so pair it with canary deployment gates.
- Gap growth: probe TTFT minus engine TTFT above a threshold, which points at the gateway.
- alert: LLMProbePathDown
expr: |
count by (probe) (
sum by (probe, job) (
increase(probe_runs_total{probe="liveness",target="public",outcome!="ok"}[2m])
) >= 3
) >= 2
labels: {severity: page}Here each external runner is a separate Prometheus job, so failures are summed per location first and the outer count is a count of locations, not of outcome series. Probes complement SLO alerts rather than replacing them; for how probe results feed an error budget see SLO engineering.
Failure modes
- Probe passes, users fail. The probe uses a short prompt and a special token with a high rate limit, while users hit long contexts or a different model alias. Mirror the real request shapes and route probes through the same auth and alias resolution.
- Cached measurements. No nonce, so the long-prefill probe reports cache hits forever.
- Flapping correctness checks. Exact string matching on sampled or batch-dependent output.
- Probe storms. Many runners, no jitter, all hitting at the same second and creating the queue they report.
- Polluted SLOs. Probe traffic counted as user traffic, masking an outage on a quiet model.
- Stale targets. The in-cluster runner keeps probing pods that were replaced, raising false alerts, or never learns about new ones, missing a bad node.
Trade-offs
| Choice | Gain | Cost |
|---|---|---|
| Higher frequency | faster detection | GPU slots, noisier SLO data |
| More locations | path coverage, quorum | runner maintenance |
| Direct-to-replica | names the bad pod | needs cluster access and discovery |
| Strict checks | catch subtle quality loss | more false alarms |
| Real user-like prompts | realistic latency | higher token cost |
What to do next
- Write one correctness probe with a computed expected answer and run it every minute from outside your network.
- Add a nonce at the start of every probe prompt and confirm the long-prefill TTFT rises accordingly.
- Tag probes with a header and exclude them from user SLO counters at the gateway.
- Add an in-cluster runner that probes each replica directly and cordons outliers automatically.
- Graph probe TTFT next to engine TTFT and investigate any persistent gap.
- Page only on location quorum and repeated wrong answers; send single-replica failures to automation.