Every distinct combination of metric name and label values is a separate time series, and every series costs memory, index space, query time and, on most hosted platforms, money. Cardinality problems rarely arrive as one bad decision. They accumulate: a label added for a debugging session, a URL path that used to be templated, a new customer tier that doubles a label's values. Then one deploy adds a user ID and the monitoring system falls over at the moment you need it most.

Managing cardinality is therefore a programme, not a cleanup. Metric cardinality, the silent cost driver explains why the cost exists. This article is about running the programme: measuring series by metric and label, attributing them to owners, setting budgets, enforcing limits at each hop of the pipeline, and redesigning the labels that explode. Examples use Prometheus and PromQL; the ideas carry over to any label-based metrics backend.

Advertisement

The arithmetic

The number of series for one metric is the number of distinct label combinations actually observed, which is bounded above by the product of each label's distinct values. Take a request-latency histogram:

LabelDistinct valuesRunning product
service4040
route (templated)25 per service, on average1,000
method4, but only 2 per route in practice2,000
status_code6 seen per route12,000
le (histogram buckets, plus +Inf)12144,000
pod8 replicas per service1,152,000

A histogram also emits _sum and _count series per combination, so the true figure is slightly higher. Two lessons follow. First, histogram buckets and the pod label are multipliers that apply to everything else; trimming 12 buckets to 8 removes a third of the series without touching any business label. Second, pod churn is cardinality over time: every rollout creates new pod names, and the old series stay in the database until they age out of retention. A service deploying ten times a day with 8 replicas creates 80 new pod values daily, which is why the active series count (in memory now) and the total stored over a retention window differ so much.

Now add user_id with 200,000 active users. Even if each user touches only a few routes, the series count jumps by orders of magnitude. No amount of hardware fixes an unbounded label; only design does.

Measure: where are the series

Prometheus exposes head-block statistics at /api/v1/status/tsdb: the top metrics by series count, the labels with the most distinct values, and the label pairs that occur in the most series. The same page is in the web UI under Status. Start there, then use PromQL for detail. Counting queries touch every series they match, so run them against a replica or at quiet times on a large server.

# Top 20 metric names by active series
topk(20, count by (__name__) ({__name__=~".+"}))

# Which label inflates one metric: distinct values per label
count(count by (route) (http_request_duration_seconds_bucket))
count(count by (pod)   (http_request_duration_seconds_bucket))

# Series each target contributes, and how many it added in the last scrape
topk(10, scrape_samples_post_metric_relabeling)
topk(10, scrape_series_added)

# The server's own total: alert on growth rate, not only the level
prometheus_tsdb_head_series
deriv(prometheus_tsdb_head_series[1h])

A spike in scrape_series_added for one job is the earliest signal of a new explosion, usually minutes after the deploy that caused it. Do not rely on rules of thumb for memory per series; they vary with label sizes, churn and version. Measure your own ratio by dividing process_resident_memory_bytes by prometheus_tsdb_head_series over a week and use that figure in capacity plans.

Advertisement

Attribute and budget

A measurement nobody owns changes nothing. Map every metric to a team: by naming prefix, by the job label, or by an ownership file in the repository that defines the metric. Then give each team a series budget, expressed the way your bill or capacity is expressed. Budgets turn a site-wide outage into a team-level conversation weeks earlier.

Cardinality controls at every hop, cheapest firstCode reviewlabel lint, budgetsSDKviews drop attributesCollector / agentdrop, aggregateScrape configrelabel, limitsTSDB / backendtenant series capsEarlier hops are cheaper to fix and more precise; later hops are blunter but cannot be bypassed.Measurement loopTSDB status APIseries by metric, labelAttributionmetric to owning teamBudget checkalert on growth, not totalOwner fixes the labelbucket, move to logs, dropnext measurementSeries = product of distinct label values actually observed, per metric, per target.One unbounded label (user id, URL path, error text) multiplies every other label.
Controls at each hop, from code review to backend caps, plus the measure-attribute-budget loop that drives them.

A workable starting point: set each team's budget at its current usage plus 20%, alert the team (not the platform on-call) when usage crosses 80% of budget, and review budgets quarterly. Alert on growth as well as level, since an explosion that will exhaust memory in three hours matters before it crosses any static threshold.

Make the budget report boring and automatic. A scheduled job can run the per-metric count query, join the result with the ownership mapping, and publish one table per team: series now, series a month ago, budget and the three metrics that grew most. Teams fix what they can see, and a monthly trend line does more than any policy document. Keep the report's own queries cheap by reading from recording rules that pre-compute the per-metric counts every few minutes, rather than running a full-index count on demand whenever someone opens the dashboard.

Enforce at every hop

No single control is enough. Earlier hops are precise and cheap to change; later hops are blunt but cannot be skipped by a team that forgot the rules.

  1. Code review and linting. Reject label names that are known to be unbounded (user_id, email, request_id, trace_id, url, raw path, error_message) in metric definitions. Require templated routes, such as /orders/{id}, from the HTTP framework's router, never the raw request path.
  2. SDK views. OpenTelemetry metric SDKs let you configure views that keep only an allow-list of attribute keys for an instrument, so an attribute added by a library never becomes a label. This is the most precise control because it runs before aggregation.
  3. Collector or agent. A pipeline stage between applications and storage can drop metrics or labels across many services at once, and is the place for organisation-wide rules.
  4. Scrape configuration. In Prometheus, metric_relabel_configs runs after the scrape and before storage, with labeldrop to remove a label and drop to discard whole series.
  5. Hard limits. Per scrape job, sample_limit caps samples per scrape after relabelling, and label_limit, label_name_length_limit and label_value_length_limit cap label count and size. Exceeding any of them fails the whole scrape: the target goes down in up and you lose all of its metrics, not just the excess.
  6. Backend tenant caps. Horizontally scaled backends such as Mimir and Cortex enforce per-tenant limits on active series and reject writes above them. Prometheus at scale covers that tier.
scrape_configs:
  - job_name: checkout
    sample_limit: 50000            # whole scrape fails above this
    label_limit: 30
    label_value_length_limit: 200  # catches stack traces and URLs in labels
    metric_relabel_configs:
      # remove a label a library adds that nobody queries
      - action: labeldrop
        regex: instance_uuid
      # discard a debug metric family entirely
      - source_labels: [__name__]
        regex: checkout_debug_.*
        action: drop

The all-or-nothing behaviour of sample_limit is a deliberate trade-off. It protects the server, but it turns a cardinality bug into a visibility outage for that target. Set limits at two to three times the target's normal output so ordinary growth never trips them, and alert on up == 0 together with scrape_samples_post_metric_relabeling so the on-call engineer can tell a crashed service from a limited one.

Redesign the labels that explode

Dropping a label removes cost and information together. Usually the information can live somewhere cheaper.

Unbounded labelWhat people wantedCheaper design
user_id, customer_idDebug one customer's latencyLogs or traces keyed by ID; metrics labelled by tier or plan
Raw URL pathLatency per endpointRouter template as the label
Error message textWhich error is risingA small error_class enum; full text in logs
trace_idJump from a spike to a traceExemplars attached to histogram samples
Exact value (bytes, item count)Distribution of sizesA histogram, not a label
Pod name on business metricsPer-replica detailAggregate with recording rules, keep pod only on runtime metrics

Exemplars deserve a special mention: they attach a trace ID to individual samples without creating a series per ID, so a dashboard can link a latency spike to a concrete trace. Span metrics shows the opposite direction, deriving low-cardinality metrics from traces. For long retention, downsampling and aggregating away the pod label after a few days cut storage further; see metric downsampling.

Histograms need their own review. Classic histograms multiply series by bucket count, so choose buckets around your SLO thresholds rather than accepting library defaults. Native histograms, where your Prometheus version and client libraries support them, store a whole distribution in one series and remove the bucket multiplier; check support across your whole pipeline before relying on them.

Worked example: a label that slipped through

A payments team ships a change that labels payment_attempts_total with merchant_id so they can see per-merchant failure rates. The figures are illustrative.

  1. Detection. Within ten minutes, scrape_series_added for the payments job rises from near zero to 15,000 per scrape, and the growth alert on deriv(prometheus_tsdb_head_series[1h]) fires to the payments team. Head series climb from 2.1 million to 2.9 million in an hour.
  2. Diagnosis. The TSDB status page lists merchant_id with 48,000 distinct values, all from one metric.
  3. Containment. A labeldrop rule for merchant_id on the payments job goes out through config, not a code deploy. New series stop at once; the existing ones age out of the head block.
  4. Redesign. The team actually needed the top merchants by failure rate. That becomes a log query on the payment-attempt events, plus a metric labelled by merchant_tier (four values) for alerting.
  5. Prevention. merchant_id joins the lint deny-list, and the job gets a sample_limit at three times its normal output.

The same incident without the growth alert usually ends with the server out of memory and every team's dashboards blank, which is why measuring growth matters more than measuring totals.

Failure modes and trade-offs

  • Over-tightening. Aggressive limits and deny-lists push teams to hide dimensions in metric names (payments_eu_failures_total), which is the same cardinality with worse query ergonomics.
  • Dropping what alerts need. A labeldrop can silently break an alert that groups by that label, which then aggregates differently or stops matching. Search alert and recording rules before dropping.
  • Duplicate series after labeldrop. If dropping a label makes two series identical, Prometheus rejects the colliding samples and counts them in prometheus_target_scrapes_sample_duplicate_timestamp_total, so values silently go missing. Drop labels only when the remaining labels still identify each series, or aggregate first.
  • Counting only active series. Churn-heavy workloads look fine on head series and expensive in long-term storage and query fan-out over weeks.
  • Platform team as the only owner. If one team pays for and polices everyone's metrics, nobody else has a reason to care. Attribution and budgets move the incentive to where labels are added.

What to do next

  1. Open the TSDB status page and record the top 20 metrics, top labels by distinct values and total head series as a baseline.
  2. Map each top metric to an owning team and publish the list.
  3. Add alerts on head-series growth rate and on per-job scrape_series_added spikes, routed to owning teams.
  4. Set sample_limit and label length limits on every scrape job at two to three times current output.
  5. Add an unbounded-label deny-list to code review and, where you use OpenTelemetry, attribute allow-lists in SDK views.
  6. Review your three largest histograms and reduce their buckets to the ones your SLOs actually use.
  7. Move any per-entity label (user, merchant, request) to logs, traces or exemplars.
  8. Review budgets and the top-20 list each quarter.
Key takeaway: Cardinality is the product of label values actually observed, multiplied by histogram buckets and replica churn, and one unbounded label can multiply everything else. Treat it as a programme: measure series by metric and label with the TSDB status API and PromQL, attribute every metric to a team, set budgets and alert on growth rather than totals. Enforce at every hop, from code review and SDK views through relabelling and scrape limits to backend tenant caps, knowing that scrape limits fail the whole target. Move per-entity detail to logs, traces and exemplars instead of labels.