An LLM observability platform records what your model-backed application actually did: which prompt version ran, what the model was sent and returned, which tools and retrievals fired, how many tokens it cost, how long each step took, and how good the answer was. Traditional APM tells you a request was slow. An LLM platform has to tell you why the answer was wrong, and that means storing the conversation itself.
That last point is what makes this a security topic. The moment you turn on tracing, the observability backend becomes the most complete copy of user input, retrieved documents and model output anywhere in your company, usually with weaker access control than the production database. This article explains the platform from the data model up: what it stores, how data flows from SDK to storage, where to redact and sample, how the evaluation loop works, how to secure the platform itself, and how to choose between hosted, self-hosted, APM-integrated and proxy-based products. The span-level detection angle is covered separately in Agent Observability, in depth.
The data model
Every product in this space uses slightly different nouns, but the data model converges on the same shapes. A trace is one end-to-end request. It contains spans (some products call them observations or runs): one per model call, tool call, retrieval, guardrail check or custom step, nested as a tree. A generation span is a model call and carries the model name, parameters, token usage, latency and, optionally, the input messages and output. Traces are grouped into sessions (a multi-turn conversation) and attributed to a user. Scores attach a number or label to a trace or span: a thumbs-down from the user, a judge model's faithfulness rating, a regex check. A prompt version links the span to the template that produced it, and datasets are curated lists of inputs, often harvested from production traces, used for offline regression runs.
OpenTelemetry's GenAI semantic conventions standardise the span part of this. The core attributes are gen_ai.operation.name (for example chat), gen_ai.provider.name, gen_ai.request.model, gen_ai.response.model, gen_ai.usage.input_tokens and gen_ai.usage.output_tokens. The conventions were still marked Development in 2026, and names have moved between releases (the provider attribute used to be gen_ai.system), so pin the convention version your instrumentation emits and map old names at the collector rather than in every dashboard. Scores, prompt versions and datasets are outside OpenTelemetry; every platform models them its own way, and that is where lock-in lives.
Architecture: data plane and control plane
The data plane has four stages. The SDK in your application creates spans, ideally through OpenTelemetry so the destination is swappable. A collector that you run receives OTLP, applies redaction, hashing, size limits and sampling, and forwards. An ingest queue absorbs bursts so a slow backend never adds latency to user requests; the SDK exports asynchronously and drops spans rather than blocking. Writers split each span into small structured metadata (ids, model, tokens, latency, status, scores) that goes into a columnar store for aggregation, and large payloads (messages, retrieved chunks, outputs) that go into a separately encrypted store with shorter retention.
The control plane is what distinguishes an LLM platform from generic tracing: a prompt registry the application fetches templates from, evaluation workers that read sampled traces and write scores, dataset curation, and alert rules over cost, error rate and score drift. Each control-plane feature reads payloads, so each is a new consumer of sensitive data that needs its own permission.
Instrumenting a model call
Instrument at the boundary where you call the model, not deep inside a framework, so the span describes what actually crossed the wire. The sketch below uses the OpenTelemetry Python API directly. Message content is attached only when an explicit flag is set, which matches the convention's own guidance that content capture is opt-in.
import os, time
from opentelemetry import trace
tracer = trace.get_tracer("support-bot")
CAPTURE_CONTENT = os.environ.get("GENAI_CAPTURE_CONTENT") == "1"
def call_model(client, model, messages, prompt_version, user_hash):
with tracer.start_as_current_span(f"chat {model}") as span:
span.set_attribute("gen_ai.operation.name", "chat")
span.set_attribute("gen_ai.provider.name", "openai")
span.set_attribute("gen_ai.request.model", model)
span.set_attribute("app.prompt.version", prompt_version)
span.set_attribute("app.user.hash", user_hash) # never the raw id
if CAPTURE_CONTENT:
span.set_attribute("app.input.messages", serialize(messages))
t0 = time.monotonic()
resp = client.chat.completions.create(model=model, messages=messages)
span.set_attribute("gen_ai.response.model", resp.model)
span.set_attribute("gen_ai.usage.input_tokens", resp.usage.prompt_tokens)
span.set_attribute("gen_ai.usage.output_tokens", resp.usage.completion_tokens)
span.set_attribute("app.latency_ms", int((time.monotonic() - t0) * 1000))
if CAPTURE_CONTENT:
span.set_attribute("app.output.text", resp.choices[0].message.content or "")
return respThree details matter. Record the response model, because aliases resolve to dated snapshots and silent upgrades cause regressions. Record the prompt version, or you cannot tell a template edit from an input shift. And hash user identifiers with a keyed hash before they leave the process, so the platform can group by user without becoming a directory of who asked what.
The collector as the policy point
The collector is the one place where every span passes through code you own, which makes it the right place for policy. Redaction belongs here rather than in each service, because services drift and a single missed call path leaks. The processor below shows the order of operations: enforce size, redact, then decide sampling with knowledge of the whole trace.
import re
PATTERNS = [
(re.compile(r"\b[\w.+-]+@[\w-]+\.[\w.]+\b"), "[EMAIL]"),
(re.compile(r"\b(?:\d[ -]?){13,19}\b"), "[CARD]"),
(re.compile(r"\bsk-[A-Za-z0-9_-]{20,}\b"), "[API_KEY]"),
]
MAX_PAYLOAD = 32_000 # characters per attribute
def redact(value: str) -> str:
value = value[:MAX_PAYLOAD]
for pattern, token in PATTERNS:
value = pattern.sub(token, value)
return value
def keep_trace(spans) -> bool:
# Tail decision: runs after the whole trace has arrived.
if any(s.status == "ERROR" for s in spans):
return True
if any(s.attrs.get("app.guardrail.flagged") for s in spans):
return True
if any(s.attrs.get("app.user.feedback") == "negative" for s in spans):
return True
return stable_hash(spans[0].trace_id) % 100 < 10 # 10% of the rest
def process(trace_spans):
for s in trace_spans:
for key in ("app.input.messages", "app.output.text", "app.retrieval.chunks"):
if key in s.attrs:
s.attrs[key] = redact(s.attrs[key])
if not keep_trace(trace_spans):
for s in trace_spans: # keep metrics, drop content
for key in [k for k in s.attrs if k.startswith("app.") and k.endswith(("messages", "text", "chunks"))]:
del s.attrs[key]
return trace_spansNotice the sampling policy never drops a span outright: it drops content. Token and latency metadata is cheap and you want it for every request, because cost dashboards built on a 10 percent sample are wrong by exactly the amount that matters at month end. Content is expensive and dangerous, so it is kept for errors, flagged traces, negative feedback and a deterministic sample keyed on trace id, so that all spans of a kept trace keep their content together. Regex redaction is a floor, not a guarantee; names and addresses in free text need an NER pass, and the patterns that fail are covered in LLM PII leakage. For the sampling machinery itself, see adaptive sampling.
The evaluation loop
Traces become valuable when something judges them. Online evaluation runs scorers over a sample of live traces and writes the result back as a score on the trace; offline evaluation replays a fixed dataset against a candidate prompt or model before release. The two share scorers, which is the point: the same faithfulness check that gates a release also watches production for drift.
def eval_worker(queue, store, judge, scorers):
for trace_id in queue: # sampled ids from the collector
t = store.load(trace_id, include_payload=True)
if t is None: # payload already expired
continue
for name, scorer in scorers.items():
if store.has_score(trace_id, name): # idempotent on redelivery
continue
result = scorer(t, judge) # {"value": 0..1, "reason": str}
store.add_score(trace_id, name, result["value"],
comment=result["reason"][:500],
scorer_version=scorer.version)
def faithfulness(t, judge):
ctx = "\n\n".join(t.retrieved_chunks)
verdict = judge(f"Context:\n{ctx}\n\nAnswer:\n{t.output}\n\n"
"Is every claim in the answer supported by the context? "
"Reply JSON {\"value\": 0 or 1, \"reason\": \"...\"}")
return parse_json(verdict)
faithfulness.version = "2026-10-01"Version every scorer, because a changed judge prompt shifts every score after it and looks exactly like a model regression on a dashboard. Make the worker idempotent, because queues redeliver. And remember that the judge is a third party receiving your production payloads: if the judge is a hosted model, it is another data processor in your privacy notice. Treat judged scores as noisy signals for trending, not verdicts on individual traces; the regression harness article covers calibrating them against human labels.
Securing the platform itself
Threat-model the platform as a database of sensitive text with a web UI, because that is what it is. The risks that come up in practice:
- Over-broad read access. Engineers get platform access to debug latency and can read every customer conversation. Separate roles for metrics and for payloads, scope projects per product and environment, and log every payload view to an audit trail.
- Key confusion. SDKs ship with keys. Use write-only ingest keys in applications and keep read or admin keys out of every deployed artefact; a leaked read key exposes history, a leaked write key only lets someone add noise.
- Stored injection in the UI. Model outputs and user inputs are attacker-controlled strings. A trace viewer that renders Markdown or HTML from them is a stored XSS vector against your own staff, and a judge that reads them is a prompt-injection target whose scores can be steered. Prompt injection forensics shows how to investigate such traces safely.
- The prompt registry as a deploy path. If the application fetches its system prompt from the platform at runtime, anyone who can edit a prompt in the UI can change production behaviour without a code review. Require review and labels such as production for promoted versions, and cache the last known good version so a platform outage does not take the application down.
- Retention and erasure. A deletion request that removes a user from the product database but not from traces, eval datasets and exported judge logs is not complete. Key everything by the hashed user id so erasure can find it, and set short payload retention by default.
Choosing a platform
Products fall into four architectural families, and the family matters more than the feature list. Licences and deployment options in this market change, so check each vendor's current terms before committing.
| Family | Examples | Strength | Watch for |
|---|---|---|---|
| Hosted LLM platform | LangSmith, Langfuse Cloud | Fast start; strong prompt, dataset and eval tooling | Payloads leave your boundary; check region, retention and sub-processors |
| Self-hosted / open source | Langfuse, Arize Phoenix | Data stays in your network; inspectable code | You run the storage tier and upgrades; check the licence terms (Phoenix uses a source-available licence) |
| APM-integrated | Datadog LLM Observability | LLM spans beside service traces, one on-call tool | Per-span pricing at LLM payload sizes; weaker dataset workflows |
| Gateway / proxy | Helicone and similar proxies | Zero-code capture of every call; caching and rate limits | Sees only model calls, not tools or retrieval; proxy is on the request path |
A sound default for most teams is OpenTelemetry instrumentation plus a collector you own, pointed at whichever backend you choose: redaction and sampling stay yours, and a later migration is a configuration change.
Worked example: sizing and exposure
Take a support assistant serving 2 million requests a day. Each trace has six spans: a router call, a retrieval, two tool calls, the main generation and an output guardrail. Average payload is 3 KB of input (system prompt, history, retrieved chunks) and 1 KB of output per generation, and span metadata is about 1 KB.
Metadata is 12 million spans a day at roughly 1 KB, about 12 GB a day before compression, kept for 90 days. Columnar compression of repetitive attributes typically shrinks this several-fold, so a few hundred gigabytes on disk is realistic. Payloads at full capture would be 2 million times about 4 KB for the main generation alone, 8 GB a day, plus retrieval chunks. With the tail policy above, suppose 1.5 percent of traces error, 2 percent are flagged and 1 percent get negative feedback; with overlap and the 10 percent sample of the rest, roughly 14 percent of traces keep content, about 1.1 GB a day. With 14 day payload retention that is around 16 GB of the most sensitive data you hold, instead of 720 GB at full capture with 90 day retention. The exposure shrank by more than an order of magnitude and you lost none of the traces you would actually open.
Online evaluation on 1 percent of traffic is 20,000 judge calls, or about 80 million judge input tokens, a day: price it before switching it on.
Failure modes
- Tracing on the critical path. A synchronous exporter or an inline proxy turns a platform outage into a product outage. Export asynchronously with bounded queues and drop on overflow; for proxies, have a bypass.
- Head sampling hides the incidents. Deciding at the first span throws away the error that happens at the fifth. Sample content at the tail, after the trace is complete.
- Lost token accounting on streams. Streaming responses often report usage only in the final chunk or not at all unless requested; spans closed early record zero tokens and cost dashboards read low.
- Unversioned prompts and scorers. Without both versions on every trace, you cannot tell a regression in the product from a change in the ruler.
- Platform as shadow production. Datasets copied from traces outlive every retention policy; give them the same classification and expiry as their source traces.
What to do next
- Inventory what you capture today: list every attribute that can hold user text and where it is stored, for how long, and who can read it.
- Move instrumentation to OpenTelemetry GenAI attributes, pin the convention version and add prompt version, response model and a keyed user hash to every generation span.
- Stand up a collector you own; put redaction, size limits and tail sampling there, keeping metadata for all traces and content only for errors, flags, feedback and a fixed sample.
- Split metadata and payload retention (for example 90 and 14 days) and wire user erasure to traces, scores and datasets.
- Separate metrics and payload roles, use write-only keys in applications, audit payload views and render trace content as plain text.
- Put the prompt registry behind review, cache last known good prompts, and version every scorer and judge prompt.
- Price online evaluation before enabling it, start with deterministic scorers, and calibrate judges against a small human-labelled set.