Prometheus is easy to start and surprisingly easy to misunderstand. You point it at a few /metrics endpoints, graphs appear, and the system seems simple. Then a graph has a gap that the service never had, increase() returns 4.3 for a counter that only moves in whole numbers, an alert stops firing during a deploy, or memory doubles after a harmless-looking release. Every one of those has the same explanation: something about how a sample travels through the server.
This article follows a single sample through a single Prometheus server: from the line a service exposes, through discovery, relabelling and the scrape, into the in-memory head block and the write-ahead log, out to compacted blocks on disk, and back through the query engine. The overview of metric types, PromQL and the wider ecosystem lives in the Prometheus deep dive, and sharding and long-term storage live in Prometheus at scale. Here the goal is that you can predict what one server will do, and fix it when it does something else.
The data model in one paragraph
A time series is identified by a metric name plus a set of label pairs, for example http_requests_total{route="/checkout",code="5xx",instance="10.0.3.7:9102"}. A sample is a timestamp in milliseconds plus a float64 value. Change any label value and you have a different series, with its own memory in the head block and its own entry in the index. That one fact explains most operational behaviour: cost scales with the number of active series, not with the number of metric names or the number of requests your service handles. A counter incremented a million times per second costs the same as one incremented once, as long as its labels stay the same.
Step 1: what the target exposes
A target serves its current values as text over HTTP. Each scrape is a snapshot: the target does not remember what Prometheus has already seen, and it does not send history. Counters are cumulative since process start, which is why a missed scrape loses resolution but not data.
# HELP http_requests_total Requests handled, by route and status class.
# TYPE http_requests_total counter
http_requests_total{route="/checkout",code="2xx"} 18234
http_requests_total{route="/checkout",code="5xx"} 41
# HELP http_request_duration_seconds Request latency.
# TYPE http_request_duration_seconds histogram
http_request_duration_seconds_bucket{route="/checkout",le="0.1"} 15002
http_request_duration_seconds_bucket{route="/checkout",le="0.5"} 18100
http_request_duration_seconds_bucket{route="/checkout",le="+Inf"} 18275
http_request_duration_seconds_sum{route="/checkout"} 2311.7
http_request_duration_seconds_count{route="/checkout"} 18275In code, a client library keeps these values in memory and renders them on request. The Python client is typical:
from prometheus_client import Counter, Histogram, start_http_server
REQUESTS = Counter("http_requests", "Requests handled, by route and status class.",
["route", "code"]) # exposed as http_requests_total
LATENCY = Histogram("http_request_duration_seconds", "Request latency.", ["route"],
buckets=(0.05, 0.1, 0.25, 0.5, 1, 2.5))
def handle_checkout(req):
with LATENCY.labels(route="/checkout").time():
resp = process(req)
# status class, never the raw user id or URL: every label value is a new series
REQUESTS.labels(route="/checkout", code=f"{resp.status // 100}xx").inc()
return resp
start_http_server(9102) # serves /metrics for the scraperThe label choice in that handler is the most expensive decision in the whole pipeline. Status class gives two or five values; the raw status code gives dozens; a user id gives millions. Metric cardinality covers how to find and fix label explosions once they have happened.
Step 2: discovery and the two relabel stages
Prometheus finds targets through service discovery (Kubernetes, Consul, cloud APIs, files, DNS). Each discovered target arrives with meta labels such as __meta_kubernetes_namespace, plus __address__, __scheme__ and __metrics_path__. Then two separate relabelling stages run, and confusing them is one of the most common configuration mistakes.
relabel_configsruns on targets, before the scrape. It decides which targets are scraped and which target labels (job, instance, namespace) will be attached to every series from them. Labels starting with__are dropped after this stage, so copy anything you want to keep into a normal label here.metric_relabel_configsruns on every scraped series, after the scrape and before storage. It is where you drop expensive metrics or labels you do not want to pay for. It does not reduce the work of the scrape itself, only what is stored.
global:
scrape_interval: 15s # default is 1m
scrape_timeout: 10s # default; must not exceed the interval
evaluation_interval: 15s # default is 1m
rule_files: ["rules/*.yml"]
scrape_configs:
- job_name: checkout
sample_limit: 20000 # fail the scrape rather than ingest an explosion
kubernetes_sd_configs: [{role: pod}]
relabel_configs: # before the scrape: which targets, which labels
- source_labels: [__meta_kubernetes_pod_label_app]
regex: checkout
action: keep
- source_labels: [__meta_kubernetes_namespace]
target_label: namespace
metric_relabel_configs: # after the scrape: which series to keep
- source_labels: [__name__]
regex: 'go_gc_.*'
action: dropBy default, honor_labels is false: if a scraped series carries a label that clashes with a target label, the server's value wins and the scraped one is renamed with an exported_ prefix. sample_limit defaults to 0, meaning no limit. Setting it is the single best defence against one bad release taking down your monitoring, because the whole scrape fails and up goes to 0 instead of millions of new series being ingested.
Step 3: the scrape loop and the series it adds
Each target gets its own scrape loop, which fires every scrape_interval (default 1m, commonly set to 15s or 30s) with jittered offsets so all targets are not hit at once. A scrape that exceeds scrape_timeout (default 10s) fails. On every attempt, successful or not, Prometheus writes a few synthetic series for the target:
| Series | What it tells you |
|---|---|
up | 1 if the scrape succeeded, 0 if it failed for any reason, including the sample limit |
scrape_duration_seconds | How long the scrape took; compare with the timeout |
scrape_samples_scraped | Samples the target exposed |
scrape_samples_post_metric_relabeling | Samples left after metric_relabel_configs |
scrape_series_added | New series this scrape created; spikes signal churn |
Alert on up == 0 per job, and graph scrape_series_added after every deploy. Current releases also support native histograms, which are switched on per scrape configuration with scrape_native_histograms (default false); without it, the classic _bucket series shown above are ingested.
Step 4: staleness
When a series that was present in the previous scrape is missing from the current one, or the target disappears or fails, Prometheus writes a staleness marker for it. Queries treat the series as ended at that point. Without markers, an instant query looks back up to the lookback delta, 5 minutes by default (--query.lookback-delta), for the newest sample, so a vanished series would linger on graphs for five minutes.
Two practical consequences follow. A failed scrape makes series disappear from instant queries straight away, so an alert like rate(errors_total[5m]) > 1 can resolve during a target outage, which is why you pair it with an up alert. And series pushed with explicit timestamps, for example through federation, do not get staleness markers, so they linger for the full lookback window.
Step 5: into the TSDB
An accepted sample is written to the write-ahead log, stored in the wal directory in 128MB segments, and appended to the head block in memory. Samples for each series are compressed into chunks. The head is what makes recent data fast to query, and the WAL is what makes it survive a crash: on restart, Prometheus replays the WAL to rebuild the head, which is why restarts of a large server can take minutes.
Ingested samples are grouped into two-hour blocks. Periodically the oldest two hours of the head are cut into a persisted block: a directory with compressed chunks, an index from label pairs to series, metadata and tombstones for deletions. The background compactor then merges blocks into larger ones, spanning up to 10% of the retention time or 31 days, whichever is smaller. Retention deletes whole blocks: 15 days by default, controlled by --storage.tsdb.retention.time and --storage.tsdb.retention.size. The documentation recommends setting the size limit to at most 80 to 85% of the disk, and states that NFS, including EFS, is not supported because non-POSIX filesystems can corrupt the database.
Capacity planning uses the documented formula: disk = retention seconds x samples per second x bytes per sample, with 1 to 2 bytes per sample on average. Worked example: 500 targets exposing 2,000 series each, scraped every 15 s, ingest 1,000,000 series / 15 s, about 66,700 samples per second. Over 15 days (1,296,000 s) at 2 bytes, that is about 173 GB. Memory is driven by active series in the head, not by retention, so the same server with 30 days of retention needs twice the disk but roughly the same RAM.
Step 6: how a query sees the data
An instant query is evaluated at one timestamp. For each selected series it takes the newest sample within the lookback window, unless a staleness marker ends the series first. A range query, which is what a graph issues, is a sequence of instant queries, one per step, which is why changing the zoom on a graph can change the shape of the line.
Range selectors like [5m] select samples in a left-open, right-closed interval: a sample exactly on the left boundary is excluded and one on the right boundary is included. rate() computes the per-second average increase over those samples, adjusts for counter resets, and extrapolates to the ends of the window to account for missed scrapes and misalignment. increase() is rate() times the window length, which is why it returns 4.3 for an integer counter. Neither is a bug; both are estimates.
Three rules follow. Make the range several times the scrape interval, at least four samples' worth, or one missed scrape leaves too few points. Apply rate() before sum(), never after, because summing counters across instances hides individual resets. And use irate() only for fast zoomed-in graphs, never for alerts, since it uses just the last two samples.
Worked example: from instrumentation to a tested alert
With the handler above exposing http_requests_total, a recording rule precomputes the per-route rate every evaluation interval and stores it as a new series, and an alert uses the stored series. Recording rules make dashboards cheap and keep the alert expression readable. The naming convention is level:metric:operations.
groups:
- name: checkout
rules:
- record: route_code:http_requests:rate5m
expr: sum by (route, code) (rate(http_requests_total[5m]))
- alert: CheckoutErrorRatioHigh
expr: |
sum(route_code:http_requests:rate5m{route="/checkout",code="5xx"})
/ sum(route_code:http_requests:rate5m{route="/checkout"}) > 0.02
for: 10m
labels: {severity: page}
annotations:
summary: "Checkout 5xx ratio above 2% for 10 minutes"The for: 10m clause keeps the alert in pending state until the condition has held across evaluations for ten minutes, which filters single spikes. Routing, grouping and silencing of the firing alert happen in Alertmanager, described in the alerting pipeline. Before deploying, check everything with promtool, and keep unit tests for alert rules in CI:
promtool check config prometheus.yml # syntax and references
promtool check rules rules/checkout.yml # PromQL parses, names are valid
promtool test rules tests/checkout_test.yml # unit-test alerts against synthetic series
promtool tsdb analyze /prometheus/data # which metrics and labels own the series
Watching the server itself
Prometheus exposes its own metrics, and a second, small Prometheus or your managed platform should scrape them. The series worth alerting on: prometheus_tsdb_head_series (active series, your main capacity number), prometheus_rule_group_iterations_missed_total (rule groups too slow for their interval), prometheus_rule_evaluation_failures_total, prometheus_tsdb_compactions_failed_total, prometheus_target_scrapes_exceeded_sample_limit_total, and disk free space against the retention size. A server that is silently missing rule evaluations still looks healthy on every dashboard it serves.
Failure modes
- Cardinality explosion: a new label with unbounded values multiplies head series, memory grows until the process is killed, and WAL replay on restart takes so long the server flaps. Mitigation: sample_limit, metric_relabel drops, tsdb analyze.
- Churn: pod restarts create new instance values, so short-lived series pile up in the head even if the active count looks stable.
- Gaps from slow scrapes: a target that takes longer than scrape_timeout fails every scrape, and up flips to 0.
- Alerts that resolve during outages: rate-based alerts lose their data when the target stops being scraped.
- Wrong relabel stage: dropping metrics in relabel_configs does nothing, and dropping a target label in metric_relabel_configs can merge series into duplicates.
- Unsupported storage: TSDB on NFS or EFS leads to corruption, and running out of disk stops compaction before retention can free space.
Trade-offs to accept knowingly
One Prometheus server is a single node with local storage and no replication. That makes it simple, fast and independent of the systems it monitors, and it means a lost disk loses history. The usual answer is two identical servers scraping the same targets, with alerts deduplicated by Alertmanager, and remote write when you need long retention or a global view. Pull scraping makes target health visible through up, but short-lived batch jobs that finish between scrapes need the Pushgateway or a different design. Accept these limits deliberately rather than discovering them during an incident.
What to do next
- Graph prometheus_tsdb_head_series and scrape_series_added for every job, and note this week's baseline.
- Set sample_limit on every scrape job at about twice its current sample count.
- Move any metric drops from relabel_configs to metric_relabel_configs, and check the opposite case too.
- Pair every rate-based alert with an up == 0 alert for the same job.
- Check that every rate() range covers at least four scrape intervals, and that rate() runs before sum().
- Run promtool check config, check rules and test rules in CI.
- Recompute disk with the documented formula, set retention.size to 80% of the volume, and confirm the volume is not NFS.