A Java service that exposes Prometheus metrics does surprisingly little work. It keeps a set of counters, gauges and histograms in memory, and when the Prometheus server sends an HTTP GET it writes their current values as text. Everything else, from storage to graphs to alerting, happens outside the JVM. That split explains most of the design rules in this article: instrumentation must be cheap on the hot path, the set of time series must stay bounded, and the values must make sense when sampled every fifteen or thirty seconds by a process that may miss a scrape.

This page is about the Java side. The Prometheus deep dive covers the server: storage, scrape scheduling and the query engine. Here we compare the three ways to get metrics out of a JVM, instrument code with the official client library and with Micrometer, pick the right metric type, keep labels under control, choose which JVM metrics deserve alerts, and walk through a latency regression that turned out to be garbage collection.

Advertisement

The pull model from the JVM's side

Prometheus pulls. The server holds a list of targets, and on each scrape interval it fetches the target's metrics endpoint, parses every line into a sample with a timestamp, and appends it to its time series database. The application never pushes and never stores history. A counter in your code is just an atomic number that only goes up; the server turns successive samples into rates with rate(). Because the server computes the rate, a process restart that resets the counter to zero is handled for you: PromQL detects the drop and treats it as a reset.

Two consequences matter for Java code. First, the cost you control is the cost of updating values, which happens on every request, not the cost of reading them, which happens a few times a minute. Second, every distinct combination of metric name and label values is a separate time series that the server stores forever, or at least until retention expires. A label holding a user id turns one metric into millions of series.

JVM processApp codecounters, timersJVM collectorsheap, GC, threadsRegistryin-memory current valuesinc / observeread MXBeans/metrics endpointHTTP server or ActuatorserializePrometheus serverscrapes every 15 to 60 sGET /metricsTSDBstores samplesRulesalerts, recordingsThe application only keeps current values; history lives in the server that pulls them
Pull-based monitoring of one JVM: application code and JVM collectors update an in-process registry, an HTTP endpoint serializes it on demand, and Prometheus scrapes it, stores the samples and evaluates alerting rules.

Three ways to expose metrics from a JVM

There are three common ways to expose metrics from a JVM, and many production services use two of them together.

ApproachHow it worksBest forWatch out for
Prometheus client_java 1.xLibrary with its own registry, metric builders and an embedded HTTP serverPlain Java services, libraries that want Prometheus-native typesYou own naming, units and the endpoint
MicrometerVendor-neutral facade; a Prometheus registry renders the outputSpring Boot, Quarkus, Micronaut; code that may export elsewhere laterNames are converted, so the exposed name differs from the code
JMX exporter agentJava agent reads MBeans and maps them to metrics with regex rulesSoftware you cannot change: Kafka, Cassandra, Tomcat, old appsRules are easy to get wrong and can explode cardinality

A sensible default: use the framework's path if you have a framework, client_java if you do not, and the JMX exporter only for third-party processes. Do not run two registries in the same JVM that both expose JVM metrics, or you will get duplicate or conflicting series under slightly different names.

Advertisement

Instrumenting code with client_java 1.x

The 1.x line of client_java was a rewrite with builder-style APIs and support for native histograms. You need three artifacts: prometheus-metrics-core for the metric types, prometheus-metrics-instrumentation-jvm for JVM metrics, and prometheus-metrics-exporter-httpserver for a small embedded endpoint. Use the current version from Maven Central for all three.

import io.prometheus.metrics.core.metrics.Counter;
import io.prometheus.metrics.core.metrics.Histogram;
import io.prometheus.metrics.exporter.httpserver.HTTPServer;
import io.prometheus.metrics.instrumentation.jvm.JvmMetrics;
import io.prometheus.metrics.model.snapshots.Unit;

public final class OrderMetrics {
    static final Counter ORDERS = Counter.builder()
        .name("orders_processed_total")
        .help("Orders processed, by outcome")
        .labelNames("outcome")              // ok | rejected | failed: a closed set
        .register();

    static final Histogram LATENCY = Histogram.builder()
        .name("order_processing_duration_seconds")
        .help("Time to process one order")
        .unit(Unit.SECONDS)
        .register();

    public static void main(String[] args) throws Exception {
        JvmMetrics.builder().register();     // heap, GC, threads, classes, process
        HTTPServer server = HTTPServer.builder().port(9400).buildAndStart();
        // ... start the application
    }

    static void handle(Order o) {
        long start = System.nanoTime();
        try {
            process(o);
            ORDERS.labelValues("ok").inc();
        } catch (RejectedException e) {
            ORDERS.labelValues("rejected").inc();
        } catch (RuntimeException e) {
            ORDERS.labelValues("failed").inc();
            throw e;
        } finally {
            LATENCY.observe((System.nanoTime() - start) / 1e9);
        }
    }
}

Three habits are visible here. Metrics are static and created once; building a metric per request re-registers it and fails. Durations are recorded in seconds, the Prometheus base unit, never in milliseconds. And the label has a closed set of values decided by the code, not by the input. If the hot path calls labelValues() heavily, keep the returned child in a field instead of looking it up each time.

Spring Boot and Micrometer

In Spring Boot, add micrometer-registry-prometheus and the Actuator starter, then expose the endpoint. Boot registers JVM, process, HTTP server, connection pool and many other metrics automatically. Micrometer 1.13, which shipped with Spring Boot 3.3, moved its Prometheus registry onto client_java 1.x; older applications use the legacy simpleclient underneath.

# application.properties
management.endpoints.web.exposure.include=health,prometheus
management.metrics.tags.application=order-service
# publish histogram buckets for HTTP server latency so PromQL can compute quantiles
management.metrics.distribution.percentiles-histogram.http.server.requests=true
@Component
class OrderHandler {
    private final Timer latency;
    private final MeterRegistry registry;

    OrderHandler(MeterRegistry registry) {
        this.registry = registry;
        this.latency = Timer.builder("order.processing")
            .publishPercentileHistogram()
            .register(registry);
    }

    void handle(Order o) {
        latency.record(() -> process(o));
        registry.counter("orders.processed", "outcome", "ok").increment();
    }
}

Micrometer names use dots and are converted at export time: order.processing becomes order_processing_seconds with _count, _sum, _max and, because of the histogram flag, _bucket series. Micrometer's JVM metrics also have their own names, for example jvm_gc_pause_seconds, which differ from client_java's. Pick one source of JVM metrics per service and write dashboards against the names it actually exposes; scrape /actuator/prometheus once by hand and read it.

The JMX exporter for code you cannot change

For a JVM you cannot recompile, the JMX exporter runs as a Java agent inside the process, reads MBeans, and serves them on its own port. A YAML file of regex rules maps MBean names and attributes to metric names and labels.

# start the JVM with
#   -javaagent:/opt/jmx_prometheus_javaagent.jar=9404:/etc/jmx/config.yaml
lowercaseOutputName: true
rules:
  - pattern: 'kafka.server<type=BrokerTopicMetrics, name=(MessagesIn|BytesIn)PerSec><>Count'
    name: kafka_server_brokertopicmetrics_$1_total
    type: COUNTER
  # no catch-all rule: unmatched MBeans are dropped instead of exported

Prefer the agent to the standalone exporter that connects over remote JMX: the agent sees process and JVM metrics directly and avoids an extra network hop. Write rules that whitelist what you need. A default catch-all rule on a Kafka broker or Cassandra node can export tens of thousands of series, many of them per topic or per table.

Choosing the metric type

TypeUse forQuery withDo not use for
CounterEvents and totals: requests, errors, bytesrate(x_total[5m])Anything that can go down
GaugeCurrent state: queue depth, pool in use, temperatureThe value, max_over_time()Event counts; you will miss events between scrapes
HistogramDistributions: latency, payload sizehistogram_quantile() over summed bucketsExact percentiles of a single instance
SummaryClient-side quantiles from one instanceThe quantile series directlyAnything you need to aggregate across instances

Histograms are the right default for latency because buckets from many instances can be summed and a quantile computed over the whole fleet; quantiles from summaries cannot be averaged in any meaningful way. Classic histograms need bucket boundaries chosen up front, and each boundary is a series. Native histograms use exponential buckets chosen automatically, store the whole distribution as one series, and give better precision. client_java 1.x histograms maintain both representations by default; the native one is only transferred over the protobuf exposition format. On the server, native histograms are stable from Prometheus 3.8 but still opt-in through the scrape_native_histograms scrape setting, so check your server version before relying on them.

Labels and cardinality in Java code

Cardinality is the number of series a metric produces: the product of the distinct values of each label. A histogram with 12 buckets, 20 endpoints, 5 status codes and 3 methods is already about 3,600 series per instance before counting _sum and _count. Multiply by 50 pods and you have 180,000 series from one metric.

  • Never put user ids, order ids, session ids, email addresses, full URLs or exception messages in labels. Put them in logs or traces, and link to them with exemplars if you need the connection.
  • Use the route template, /orders/{id}, not the raw path. Spring and most frameworks provide it; if yours does not, normalise in a filter and map unknown paths to one value.
  • Collapse status codes to classes such as 2xx and 5xx unless you really alert on individual codes.
  • Do not use exception class names as labels without a fallback bucket; a new library version can introduce dozens.

Micrometer can enforce limits centrally with a MeterFilter, for example MeterFilter.maximumAllowableTags to cap the distinct values of a tag. The cardinality management guide covers finding and fixing an explosion once it has reached the server.

JVM metrics and the PromQL that reads them

JVM metrics are useful in a small set. The names below are the ones client_java's JvmMetrics exposes; adjust them if Micrometer is your source.

# Heap in use as a fraction of max (alert when sustained above about 0.9)
sum by (instance) (jvm_memory_used_bytes{area="heap"})
  / sum by (instance) (jvm_memory_max_bytes{area="heap"})

# Fraction of wall-clock time spent in GC over 5 minutes
sum by (instance) (rate(jvm_gc_collection_seconds_sum[5m]))

# Thread count growth, a leak signal
deriv(jvm_threads_current[30m])

# CPU used by the process, in cores
rate(process_cpu_seconds_total[5m])

# p99 request latency across the fleet from a classic histogram
histogram_quantile(0.99,
  sum by (le) (rate(order_processing_duration_seconds_bucket[5m])))

Two notes. jvm_gc_collection_seconds is a summary per collector, so its rate of _sum is seconds of GC per second; above 0.1 means the application loses more than a tenth of its time to collection. That reading holds for stop-the-world collectors such as G1, Parallel and Serial; ZGC and Shenandoah also report concurrent cycle time, so filter on the gc label to the pause collector before treating the sum as time lost. And a heap ratio of 0.9 is not by itself a problem for G1, which uses the heap it is given; what matters is whether it stays high after collections and whether GC time rises with it. The G1 collector article explains how to read those patterns.

Worked example: a latency regression caused by GC

A checkout service with 12 pods began breaching its p99 latency objective of 300 ms several times an hour. Average latency barely moved. The numbers below are illustrative, but the sequence is how such incidents are usually solved.

  1. The fleet-wide p99 from the histogram rose from 180 ms to 650 ms in short spikes. Splitting the same query by (instance) showed the spikes moved between pods rather than hitting all of them at once, which rules out a shared downstream dependency.
  2. On the affected pods, rate(jvm_gc_collection_seconds_sum[1m]) spiked to about 0.4 at the same moments, while request rate and CPU stayed flat.
  3. The heap ratio after each collection had crept from 0.45 to 0.8 since a release two days earlier, and the old generation pool series showed steady growth: a cache added in that release had no size bound.
  4. The fix bounded the cache. The team then added an alert on GC time above 0.15 for ten minutes and a recording rule for p99 per instance, so the next time a single pod misbehaves the dashboard shows it immediately.

The lesson is not that GC causes latency. It is that per-instance breakdowns, a GC time rate and a post-collection heap trend together answer in minutes a question that averages hide entirely.

Failure modes and trade-offs

FailureSymptomFix
Unbounded labelsServer memory grows, scrapes slow, queries time outRemove the label, normalise values, add a filter cap
Metric created per callRegistration errors or a leak of meter objectsCreate metrics once as fields or static constants
Milliseconds in base metricsDashboards off by 1000x, mismatched bucketsRecord seconds and bytes; convert only in display
Averaging summary quantilesFleet p99 looks healthy while one pod suffersUse histograms and aggregate buckets
Slow scrape endpointGaps and up == 0 under loadKeep collectors cheap; never do I/O inside a callback
Short-lived jobsBatch finishes between scrapes; nothing recordedUse a Pushgateway only for job completion metrics, or a longer-lived process
Duplicate JVM metricsTwo registries, conflicting namesOne source of JVM metrics per process

The main trade-off is resolution against cost. Shorter scrape intervals and more buckets show more detail but multiply storage. For most services a 15 to 30 second interval, a dozen buckets or a native histogram, and fewer than a few thousand series per instance is a comfortable budget. Prometheus metrics also tell you that something is wrong, not which request; pair them with tracing, for example through OpenTelemetry for Java, when you need per-request detail.

What to do next

  1. Scrape your service's endpoint by hand with curl and list every metric name and label; delete what nobody uses.
  2. Count series per instance with count by (__name__) ({instance="host:port"}) and investigate any metric above a few hundred.
  3. Make sure latency is a histogram recorded in seconds, with route templates, not raw paths, as labels.
  4. Add the four JVM queries above to a dashboard: heap after GC, GC time rate, thread count trend and process CPU.
  5. Create alerts on symptoms users feel, error rate and p99 latency, and use JVM metrics for diagnosis rather than paging.
  6. Decide whether your server will scrape native histograms, and if so enable scrape_native_histograms for the job.
  7. For third-party JVMs, replace any catch-all JMX exporter rule with an explicit whitelist.
Key takeaway: A Java service only keeps current values in memory and lets Prometheus pull them, so the job in code is cheap updates and a bounded set of series. Use the framework path or client_java for your own code and the JMX exporter only for software you cannot change. Record latency as a histogram in seconds, keep labels to small closed sets, choose one source of JVM metrics, and alert on user-facing symptoms while using heap, GC time and thread trends to diagnose.