Logs are the oldest observability signal and usually the most expensive. Every service writes them, few teams agree on their shape, and the bill grows with traffic, verbosity and retention at once. During an incident they are often the only record of what one request actually did, which is exactly when a weak pipeline turns out to be dropping lines, missing the trace ID, or taking minutes to answer a query.
This article treats logging as a data pipeline and explains each stage from first principles: what goes into an event, how lines leave the node, where to parse, redact and sample, how buffers behave under back-pressure, how the main storage designs index data, and how retention drives cost. Choosing a backend is covered in Loki vs Elastic vs ClickHouse for logs; here the focus is the architecture around whichever one you pick.
What logs are for, and what they are not
Metrics are pre-aggregated numbers that tell you that something is wrong. Traces record a request's path across services and tell you where time went. Logs are discrete events with arbitrary detail and tell you what exactly happened: which order failed, which rule rejected it, what the downstream error said.
That detail is the cost. Each line can carry high-cardinality values and is stored individually rather than summed. So if you count log lines to draw a graph, that number should probably be a metric emitted by the service, with the log line kept as evidence for when the graph moves.
The pipeline end to end
The application formats an event and writes it, in containers to stdout. The node (container runtime and kubelet on Kubernetes) captures that stream into files under /var/log/pods and rotates them. A shipper such as Fluent Bit, Vector or the OpenTelemetry Collector's filelog receiver tails the files, remembers its position, and parses and enriches each line. A buffer absorbs bursts and outages: the shipper's disk buffer or a Kafka topic. The store indexes and compresses, and the query layer serves searches, dashboards and alerts, with an object-storage archive alongside for long retention.
Two properties matter more than any product choice. The pipeline is asynchronous, so the application must never wait for the store and every stage needs a policy for a slow next stage. And it is lossy by default: rotation, full buffers and rejected writes drop data silently. Decide where back-pressure stops and where data is dropped, and export a counter for every drop.
Emitting: structured events and a schema
Write one JSON object per line, not prose to grep. {"level":"INFO","msg":"payment declined","reason":"insufficient_funds"} can be filtered, grouped and joined; the same fact in a sentence needs a regular expression that breaks when someone rewords it.
Agree a small mandatory schema. The OpenTelemetry log data model is a good reference: timestamp and observed timestamp, severity text and number, body, trace ID, span ID and trace flags, resource attributes describing the service, and free-form attributes. Every event should carry a UTC millisecond timestamp, level, service, environment, the trace and span IDs when a span is active, and a stable message; use semantic conventions for shared attribute names. Keep the message constant and put variable parts in fields so events can be counted. The Python example formats JSON, attaches the active trace context, and uses a queue so the request thread never blocks on I/O.
import json, logging, logging.handlers, os, queue, sys, time
from opentelemetry import trace
SERVICE = os.environ.get("OTEL_SERVICE_NAME", "checkout")
class JsonFormatter(logging.Formatter):
def format(self, record):
event = {
"ts": time.strftime("%Y-%m-%dT%H:%M:%S", time.gmtime(record.created))
+ ".%03dZ" % record.msecs,
"level": record.levelname,
"service": SERVICE,
"logger": record.name,
"msg": record.getMessage(), # a constant message, values go in fields
}
ctx = trace.get_current_span().get_span_context()
if ctx.is_valid: # join key to traces
event["trace_id"] = format(ctx.trace_id, "032x")
event["span_id"] = format(ctx.span_id, "016x")
event.update(getattr(record, "fields", {}))
if record.exc_info:
event["exception"] = self.formatException(record.exc_info)
return json.dumps(event, default=str)
# The request thread only enqueues; a background listener does the write.
q = queue.Queue(maxsize=10000)
stdout = logging.StreamHandler(sys.stdout)
stdout.setFormatter(JsonFormatter())
listener = logging.handlers.QueueListener(q, stdout)
root = logging.getLogger()
root.addHandler(logging.handlers.QueueHandler(q))
root.setLevel(logging.INFO)
listener.start()
log = logging.getLogger("payments")
log.info("payment declined",
extra={"fields": {"order_id": "o-81723", "reason": "insufficient_funds",
"amount_minor": 4599, "currency": "EUR"}})The bounded queue is deliberate. If stdout blocks, for example on a full node disk, QueueHandler drops records through handleError rather than stalling requests. Audit logs, where loss is unacceptable, belong on a separate synchronous and durable path.
Collecting on the node
On Kubernetes the runtime writes each container's output to a file, prefixing lines with a timestamp, stream name and partial-line flag, and the kubelet rotates the file at a size limit. The shipper, running as a DaemonSet, tails every file, rejoins partial lines and adds namespace, pod and label metadata from the API server.
It records a checkpoint, a file identity and byte offset, so a restart resumes rather than re-reading or skipping. This is where the first silent losses happen. If the shipper lags and the kubelet rotates and deletes a file before it is fully read, those lines are gone; if the checkpoint sits on an ephemeral volume, a restart re-reads everything. Keep checkpoints on a host path, size rotation to hold several minutes of shipper outage, and alert on shipper lag. Emit stack traces as one JSON field, as above, to avoid fragile multiline parsing.
Processing: parse, enrich, redact and sample
The shipper is the last point before data leaves the host, and work there is spread across every node. Four jobs belong here: parse JSON to top-level fields, enrich with platform metadata, redact secrets and personal data before they are replicated and backed up, and sample or drop noise such as health checks and bulk DEBUG output. The Vector configuration below does all four and fans out to a sampled searchable store and a complete compressed archive.
sources:
pods:
type: kubernetes_logs
transforms:
parse:
type: remap
inputs: [pods]
source: |
parsed, err = parse_json(.message)
if err == null && is_object(parsed) {
. = merge(., object!(parsed))
del(.message)
}
if !exists(.service) {
.service = .kubernetes.container_name
}
del(.password)
del(.authorization)
if is_string(.msg) {
.msg = replace(string!(.msg), r'\b\d{13,19}\b', "[REDACTED_PAN]")
}
sample_debug:
type: sample
inputs: [parse]
rate: 20 # keep 1 in 20 ...
exclude: '.level != "DEBUG"' # ... but only DEBUG lines are sampled
sinks:
loki:
type: loki
inputs: [sample_debug]
endpoint: http://loki-gateway.logging:3100
encoding:
codec: json
labels: # low-cardinality only
service: "{{ service }}"
namespace: "{{ kubernetes.pod_namespace }}"
level: "{{ level }}"
buffer:
type: disk
max_size: 1073741824 # 1 GiB; the disk minimum is about 256 MiB
when_full: block
archive:
type: aws_s3
inputs: [parse] # unsampled copy for audit and replay
bucket: acme-log-archive
region: eu-west-1
key_prefix: "logs/{{ service }}/%Y/%m/%d/"
compression: gzip
encoding:
codec: jsonThe Loki label set is the most important line. Loki indexes only labels, and each distinct label combination is a separate stream: service, namespace and level give hundreds of streams, while a user or trace ID would give millions. Other stores show the same problem as field-mapping explosion, the log equivalent of metric cardinality. Keep identifiers in the body and filter at query time. Test redaction rules in CI against lines with fake card numbers, because a regex that never matches looks exactly like one that works.
Buffering and back-pressure
Backends slow down during compaction and incidents, exactly when volume peaks. Each buffer can block, pushing back upstream, or drop. Blocking in the shipper is usually right, because the node's files are a further buffer behind it, but back-pressure must stop at the node and never reach the application.
Memory buffers are lost on restart; disk buffers survive but cost IOPS, and in Vector have a minimum of roughly 256 MiB. At scale, a Kafka tier between shippers and indexers decouples collection from indexing, lets several consumers read one stream, and allows replay after an indexing bug, at the cost of another stateful system. Alert on buffer fill and dropped-event counters: a full buffer is the earliest warning of loss.
Storage, indexing and queries
Three designs decide which queries are fast. Inverted indexes in Elasticsearch and OpenSearch index every token: ad hoc search is fast, but ingestion is CPU-heavy, the index adds substantially to storage, and dynamic mappings can explode. Label-indexed chunk stores such as Loki index a few labels and keep compressed chunks in object storage: cheap, but a query scans every chunk in the selected streams, so selectors must be narrow. Columnar stores such as ClickHouse store fields as columns: aggregations are very fast and compression excellent, but you need a schema and a sort key built around common filters. All three tier data by age, through lifecycle policies, object-storage flushes or TTL moves.
Most incident queries either pull everything about one request or count how often something happens. The LogQL below does both; with trace context propagated, the first is one click from a slow span.
# every line of one request, across all services
{namespace="shop"} | json | trace_id="4bf92f3577b34da6a3ce929d0e0e4736"
# declines per reason over five-minute windows, usable as an alert or a panel
sum by (reason) (
count_over_time({service="checkout", level="INFO"} | json | msg="payment declined" [5m])
)
Worked example: sizing a pipeline
Take 400 pods emitting 20 lines per second at 600 bytes each: 8,000 lines and 4.8 MB per second, about 415 GB of raw logs a day. JSON logs typically compress 5 to 15 times; at an assumed 10, that is about 41 GB a day. Measure your own ratio first.
Thirty days searchable in Loki is then about 1.25 TB of object storage plus a small index, and a year of archive about 15 TB before cheaper storage classes. An inverted-index store needs far more fast disk for the same 30 days, since the index and usually a replica are added, so size it from a test ingest of real logs.
Now one service carrying 5 percent of traffic switches to DEBUG and emits 10 times its volume: total volume rises 45 percent in minutes. The sampler cuts its DEBUG lines to one in twenty, so the searchable store grows by just over 2 percent while the archive keeps everything.
Failure modes and trade-offs
- Silent loss at rotation. Compare emitted and stored line counts for a canary service.
- Cardinality blow-up. An ID used as a label slows everyone. Review label sets like schema changes.
- Secrets in the store. Redact at the shipper and document how to purge every tier and backup when it fails.
- Back-pressure reaching the app. A synchronous remote appender turns a logging outage into a service outage.
- Clock skew. Stores that order by timestamp reject or misplace late lines; keep clocks synced.
- Cost creep. Publish bytes per service per day and give each team a budget.
| Decision | Option A | Option B |
|---|---|---|
| Index scope | Everything: fast ad hoc search, high cost | Labels only: cheap, needs narrow selectors |
| Sampling | Keep all lines: complete evidence | Sample noisy levels: cheaper, gaps in detail |
| Buffer | Shipper disk buffer: simple | Kafka tier: replay and fan-out, one more system |
| Redaction | At the shipper: never leaves unredacted | At the store: central rules, secrets already in transit |
What to do next
- Write a mandatory log schema (timestamp, level, service, environment, trace ID, span ID, message) and ship a shared formatter for your main languages.
- Log JSON to stdout through a non-blocking handler; move audit events to a durable path.
- Put checkpoints on host paths, size rotation for shipper outages, and alert on lag, buffer fill and drops.
- Add shipper redaction rules with fake-secret test fixtures in CI.
- Keep indexed labels low-cardinality, sample DEBUG and health-check noise, and archive an unsampled copy.
- Wire trace-to-log links into dashboards and publish bytes per service per day.