Micrometer is to metrics what SLF4J is to logging: you instrument your code once against a vendor-neutral API, and a registry for your monitoring system turns those measurements into whatever that system expects. Prometheus scrapes a text page, Datadog and CloudWatch receive pushes, and OTLP goes to a collector. The calling code does not change.

This page is about the library itself: how a meter is identified, which meter type fits which question, how timers turn samples into percentiles and histograms, how filters and registries compose, and how the Observation API produces metrics and traces from one block of code. Prometheus exposition and PromQL are covered in Prometheus monitoring for Java, and OpenTelemetry in OpenTelemetry for Java. Every behaviour shown below was run against micrometer-core 1.16.2. 1.17 is the newest stable line at the time of writing, and the APIs used here are the same.

The model: meters, registries and observations

Three pieces make up the model. A meter is a named instrument, such as a counter or timer, that you record into. A MeterRegistry creates meters and holds their state; each monitoring system has its own subclass, which decides how values are stored and shipped. A CompositeMeterRegistry forwards to several registries at once, so the same meter can feed Prometheus and a push backend during a migration.

Registration passes through the registry's filter chain, and the resulting ID indexes a map of live meters. Library meters (JVM, connection pools, HTTP clients) take the same path as yours. The Observation API sits beside this: you describe an operation once, and handlers turn it into meters or spans.

Your codecounter(), timer(), observe()Instrumented librariesHTTP client, JDBC pool, JVMMeterRegistryfilters, then Meter.Id lookupCompositeMeterRegistryfans out to childrenPrometheus registrycumulative, scrapedStep registryper-interval, pushedObservationRegistryone context, many handlersTracing handlerspans via a bridgeregisteradd()meter handlerObservation: one timed block becomes a timer, a long-task timer and a span
Where a measurement goes: the registry filters and deduplicates meters, a composite fans out to backend registries, and an ObservationRegistry hands one timed block to several handlers.

Meter identity: name plus tags

A meter's identity is its name plus its full set of tags (key-value pairs) and its type. Ask the registry twice for the same identity and you get the same object back. That is why it's safe to look meters up in a hot path, although a lookup still costs a map access that a cached field avoids. Change one tag value and you have a new meter, which is a new time series in the backend.

SimpleMeterRegistry reg = new SimpleMeterRegistry();
Counter a = reg.counter("orders", "status", "ok");
Counter b = reg.counter("orders", "status", "ok");
System.out.println(a == b);                       // true: same Meter.Id

Names are lowercase and dotted, like http.client.requests, and each registry rewrites them for its backend: the Prometheus registry turns a timer payment.authorize into payment_authorize_seconds_count, _sum and _max. Don't put units or backend suffixes in names yourself.

Tag values are the cardinality budget: every distinct combination is a separate series, so tags must come from small, bounded sets such as a route template, a status class or a region. Never a user ID, order ID or raw path.

Choosing a meter type

QuestionMeterNotes
How many times did X happen?CounterMonotonic. Rates are computed in the backend, never in code.
What is X right now?GaugeSampled when read. Holds a weak reference to the observed object by default.
How long did X take, and how often?TimerCount, total time, max, and optional percentiles or histogram.
How big was X?DistributionSummaryA timer for non-time values: payload bytes, batch sizes.
How long has the in-flight X been running?LongTaskTimerShows active tasks and their duration while they run.
A library already counts XFunctionCounter, FunctionTimerRead an existing counter rather than double-counting.

Gauges are the type people get wrong most often. A gauge calls your function each time it's read, and Micrometer holds the observed object weakly to avoid leaks. If nothing else references it, it's collected and the gauge reads NaN, as happened when we gauged a list only the gauge referenced and forced a GC.

// Wrong: the list is only reachable through the gauge, so it can be collected.
Gauge.builder("queue.size", new ArrayList<>(List.of(1, 2)), List::size).register(reg);
// after GC: gauge value -> NaN

// Right: keep a strong reference in a field you own, or say so explicitly.
private final Queue<Job> queue = new ConcurrentLinkedQueue<>();
Gauge.builder("queue.size", queue, Queue::size).register(reg);
Gauge.builder("cache.entries", cache, Cache::size).strongReference(true).register(reg);

To time a span that ends in another method, use Timer.start(registry) and later sample.stop(timer); the timer is chosen at stop time, so it can carry the outcome tag.

Timers, percentiles and histograms

A plain timer publishes count, total time and max. Percentiles are opt-in, and there are two ways to get them, which differ in what you can do with them later.

Client-side percentiles (publishPercentiles(0.95)) are computed inside your JVM and shipped as ready-made numbers. They are approximations over a sliding time window. That window is the expiry, which defaults to the registry's step, multiplied by the buffer length, which defaults to 3. Recording a single 30 ms sample and reading the snapshot reported p95 as 29.36 ms: the internal histogram's resolution, not the recorded value. The bigger limitation is in the documentation: percentile approximations cannot be aggregated across tags. Averaging the p95 of ten pods does not give the fleet p95.

Percentile histograms (publishPercentileHistogram()) ship bucket counts instead, and backends that support it, such as Prometheus, compute quantiles at query time over any aggregation you want. Micrometer's bucket generator produces 276 buckets. It ships only those within the expected minimum and maximum, which is 73 per timer per tag combination by default. Narrowing that range is the main control over cost.

Timer authorize = Timer.builder("payment.authorize")
    .tag("provider", "acme")
    .publishPercentileHistogram()                     // aggregable buckets
    .serviceLevelObjectives(Duration.ofMillis(300))   // an exact bucket at your SLO
    .minimumExpectedValue(Duration.ofMillis(5))
    .maximumExpectedValue(Duration.ofSeconds(5))      // fewer buckets shipped
    .register(registry);

Use histograms for anything you alert on or aggregate, with an SLO bucket so "fraction under 300 ms" is exact rather than interpolated.

Meter filters and why their timing matters

Every registration passes through the registry's MeterFilter chain, in the order the filters were added. A filter can deny or accept a meter, rewrite its ID (rename it, add common tags, drop or replace tag values), or change its distribution configuration. Filters are where you enforce cardinality and naming policy once, for code you don't own as well as your own.

registry.config()
    .commonTags("app", "checkout", "region", region)
    .meterFilter(MeterFilter.denyNameStartsWith("jvm.gc.pause.debug"))
    .meterFilter(MeterFilter.maximumAllowableTags(
        "http.client", "uri", 100, MeterFilter.deny()))     // cap a runaway tag
    .meterFilter(new MeterFilter() {
        @Override
        public DistributionStatisticConfig configure(Meter.Id id, DistributionStatisticConfig cfg) {
            if (id.getName().startsWith("payment.")) {
                return DistributionStatisticConfig.builder()
                    .percentilesHistogram(true).build().merge(cfg);
            }
            return cfg;
        }
    });

With a cap of 2 on uri and requests to four paths, the registry kept the first two series and denied the rest. That is the intended behaviour: a cap protects the backend, and the data it drops tells you a tag needs templating.

Order matters, and timing matters more. When we added a common-tags filter after a counter called early was already registered, Micrometer logged "A MeterFilter is being configured after a Meter has been registered to this registry", and asking for early again created a second, tagged meter. The original untagged meter remained, still held by the code that had it. The result is two series under one name, each holding part of the count. Configure every filter before the first meter exists. In Spring Boot, a MeterFilter or MeterRegistryCustomizer bean runs early enough. A filter added from your own startup code may not.

Cumulative and step registries

Registries fall into two families. Cumulative registries, such as Prometheus, report totals since start, and the backend computes rates from the difference between scrapes. A missed scrape loses resolution but not data. Step registries, used by push systems like Datadog or CloudWatch, report what happened in each interval (the step, often one minute) and reset afterwards. A failed push loses that interval's data unless the registry retries.

The family changes what your numbers mean. A step registry's counter value is a per-interval count, and its percentile window is tied to the step. If you put both registry types in a composite during a migration, the backends will chart "the same" meter differently. That is expected, not a bug. Shut down cleanly: call registry.close() so push registries flush the last partial step. Otherwise the last minute before every deploy goes missing.

Avoid the static Metrics.globalRegistry in application code: recordings made before a real registry is added go nowhere, and tests share hidden state. Inject a MeterRegistry instead.

The Observation API

The Observation API, in micrometer-observation, lets you describe an operation once: its name, its context, and its key-values. Handlers you register decide what to produce from it. The metrics handler makes meters, and a tracing bridge makes spans. Key-values are split by cardinality. Low-cardinality values become metric tags as well as span attributes. High-cardinality values, such as an order ID, go to spans only, where per-request detail costs nothing extra.

ObservationRegistry obs = ObservationRegistry.create();
obs.observationConfig().observationHandler(new DefaultMeterObservationHandler(meterRegistry));

Observation.createNotStarted("payment.authorize", obs)
    .lowCardinalityKeyValue("provider", "acme")
    .highCardinalityKeyValue("order.id", orderId)
    .observe(() -> client.authorize(order));

Running exactly this against a registry with the common tag app=checkout produced two meters. The first was a TIMER named payment.authorize with tags app, error=none, provider. The second was a LONG_TASK_TIMER named payment.authorize.active with tags app, provider. order.id appeared in neither, which is the protection working. The handler adds the error tag itself: it holds the exception's class name when the block throws, so failure rates come from the same timer. Two details follow from this. Key-values added after the observation starts can't reach the .active meter, which was created at start. And in Spring Boot 3, the HTTP server and client metrics are themselves observations, so customise them through observation conventions rather than the older metrics-only hooks.

Worked example: instrumenting a payment client

Here is a payment client instrumented properly, end to end. The requirements are: a request rate by provider and outcome, an exact fraction of calls under 300 ms, in-flight calls during an incident, and traces carrying the order ID.

public final class PaymentClient {
    private final ObservationRegistry obs;
    private final HttpPaymentApi api;

    public PaymentClient(ObservationRegistry obs, HttpPaymentApi api) {
        this.obs = obs;
        this.api = api;
    }

    public AuthResult authorize(Order order) {
        return Observation.createNotStarted("payment.authorize", obs)
            .lowCardinalityKeyValue("provider", api.providerName())   // 3 values
            .highCardinalityKeyValue("order.id", order.id())          // spans only
            .observe(() -> api.authorize(order));                     // error tag on throw
    }
}

// At startup, before any meter exists:
registry.config().meterFilter(new MeterFilter() {
    @Override
    public DistributionStatisticConfig configure(Meter.Id id, DistributionStatisticConfig c) {
        if (!id.getName().equals("payment.authorize")) return c;
        return DistributionStatisticConfig.builder()
            .percentilesHistogram(true)
            .serviceLevelObjectives(Duration.ofMillis(300).toNanos())
            .maximumExpectedValue((double) Duration.ofSeconds(5).toNanos())
            .build().merge(c);
    }
});

On start the handler begins a sample and the active long-task timer; on stop the timer records the duration with error set, and the Prometheus registry exposes buckets including one at 0.3 s. Series count per instance is 3 providers times the error classes you see times the bucket count, so estimate it before shipping. The SLO is le="0.3" bucket count divided by total count, summed across pods, and that sum is valid because buckets aggregate.

Testing your metrics

Test metrics like any other output. SimpleMeterRegistry keeps everything in memory. Inject it, run the code, and assert on what was recorded. For observations, the separate micrometer-observation-test module provides a test registry with fluent assertions, which is useful when you want to check key-values without involving meters.

@Test
void failedAuthorizationIsTimedWithErrorTag() {
    SimpleMeterRegistry meters = new SimpleMeterRegistry();
    ObservationRegistry obs = ObservationRegistry.create();
    obs.observationConfig().observationHandler(new DefaultMeterObservationHandler(meters));

    var client = new PaymentClient(obs, failingApi());
    assertThrows(PaymentException.class, () -> client.authorize(order()));

    Timer t = meters.get("payment.authorize").tag("error", "PaymentException").timer();
    assertEquals(1, t.count());
}

Use a fresh registry per test, and snapshot the list of meter names your service registers: dashboards depend on them, so a rename is a breaking change.

Failure modes

  • Series explosion. A raw path, user ID or exception message used as a tag. Symptom: the backend's memory or bill climbs and scrapes slow down. Fix: template the value, or cap it with maximumAllowableTags and watch for denials.
  • Gauge reads NaN. The observed object was garbage collected. Fix: hold a strong reference you own, or use strongReference(true).
  • Split series after startup. A filter was registered after meters existed, as shown above. Fix: configure filters first, and treat the warning as an error in CI logs.
  • Averaged percentiles. A dashboard averages client-side p99 across pods. Fix: switch to histograms and compute quantiles in the backend.
  • Inconsistent tag keys. One name registered with different tag key sets from two code paths; Prometheus needs one key set per name, and depending on version the registry throws or drops the meter. Register each name in one place.
  • Lost final interval. A push registry is not closed on shutdown. Fix: close the registry, or let the framework do it.

Trade-offs

OptionChoose it whenCost
Micrometer APIJVM services, Spring Boot, several possible backends, existing library bindersOne more abstraction; its naming conventions sit between you and the backend
OpenTelemetry API directlyPolyglot fleet standardising on OTel, signals correlated in one SDKFewer ready-made JVM library binders in some stacks; different idioms
Prometheus client_java directlyA Prometheus-only shop that wants exact control of expositionYou lose backend portability and Micrometer's binders

For most JVM services Micrometer is the default because frameworks already emit through it; the real design question is how much cardinality and histogram cost each meter deserves. JFR is the right tool for per-event detail that metrics should never carry; see Java Flight Recorder.

What to do next

  1. Inject a MeterRegistry everywhere you record. Remove uses of Metrics.globalRegistry from application code.
  2. List every tag your services emit, along with its realistic number of values. Template or remove any that are unbounded, and add maximumAllowableTags as a backstop.
  3. Move filter configuration to the earliest startup hook. In Spring Boot, these are filter and customizer beans (see Spring Boot).
  4. Switch the latency meters you alert on to percentile histograms, with an SLO bucket and a narrowed expected range. Estimate the series count first.
  5. Audit gauges for weak references to objects nothing else holds.
  6. Wrap your key outbound calls in Observations, with low-cardinality outcome tags and high-cardinality IDs.
  7. Add a test that snapshots the list of meter names, and a SimpleMeterRegistry test for every meter that feeds an alert.
Key takeaway: A Micrometer meter is identified by its name and full tag set, so tag values are a cardinality budget to spend deliberately. Use histograms rather than client-side percentiles for anything you aggregate or alert on, configure every MeterFilter before the first meter exists, keep strong references behind gauges, and describe important operations as Observations so one block of code yields a timer, an active-task timer and a span. Test meters with SimpleMeterRegistry like any other output.