A Java service that exposes Prometheus metrics does surprisingly little work. It keeps a set of counters, gauges and histograms in memory, and when the Prometheus server sends an HTTP GET it writes their current values as text. Everything else, from storage to graphs to alerting, happens outside the JVM. That split explains most of the design rules in this article: instrumentation must be cheap on the hot path, the set of time series must stay bounded, and the values must make sense when sampled every fifteen or thirty seconds by a process that may miss a scrape.
This page is about the Java side. The Prometheus deep dive covers the server: storage, scrape scheduling and the query engine. Here we compare the three ways to get metrics out of a JVM, instrument code with the official client library and with Micrometer, pick the right metric type, keep labels under control, choose which JVM metrics deserve alerts, and walk through a latency regression that turned out to be garbage collection.
The pull model from the JVM's side
Prometheus pulls. The server holds a list of targets, and on each scrape interval it fetches the target's metrics endpoint, parses every line into a sample with a timestamp, and appends it to its time series database. The application never pushes and never stores history. A counter in your code is just an atomic number that only goes up; the server turns successive samples into rates with rate(). Because the server computes the rate, a process restart that resets the counter to zero is handled for you: PromQL detects the drop and treats it as a reset.
Two consequences matter for Java code. First, the cost you control is the cost of updating values, which happens on every request, not the cost of reading them, which happens a few times a minute. Second, every distinct combination of metric name and label values is a separate time series that the server stores forever, or at least until retention expires. A label holding a user id turns one metric into millions of series.
Three ways to expose metrics from a JVM
There are three common ways to expose metrics from a JVM, and many production services use two of them together.
| Approach | How it works | Best for | Watch out for |
|---|---|---|---|
| Prometheus client_java 1.x | Library with its own registry, metric builders and an embedded HTTP server | Plain Java services, libraries that want Prometheus-native types | You own naming, units and the endpoint |
| Micrometer | Vendor-neutral facade; a Prometheus registry renders the output | Spring Boot, Quarkus, Micronaut; code that may export elsewhere later | Names are converted, so the exposed name differs from the code |
| JMX exporter agent | Java agent reads MBeans and maps them to metrics with regex rules | Software you cannot change: Kafka, Cassandra, Tomcat, old apps | Rules are easy to get wrong and can explode cardinality |
A sensible default: use the framework's path if you have a framework, client_java if you do not, and the JMX exporter only for third-party processes. Do not run two registries in the same JVM that both expose JVM metrics, or you will get duplicate or conflicting series under slightly different names.
Instrumenting code with client_java 1.x
The 1.x line of client_java was a rewrite with builder-style APIs and support for native histograms. You need three artifacts: prometheus-metrics-core for the metric types, prometheus-metrics-instrumentation-jvm for JVM metrics, and prometheus-metrics-exporter-httpserver for a small embedded endpoint. Use the current version from Maven Central for all three.
import io.prometheus.metrics.core.metrics.Counter;
import io.prometheus.metrics.core.metrics.Histogram;
import io.prometheus.metrics.exporter.httpserver.HTTPServer;
import io.prometheus.metrics.instrumentation.jvm.JvmMetrics;
import io.prometheus.metrics.model.snapshots.Unit;
public final class OrderMetrics {
static final Counter ORDERS = Counter.builder()
.name("orders_processed_total")
.help("Orders processed, by outcome")
.labelNames("outcome") // ok | rejected | failed: a closed set
.register();
static final Histogram LATENCY = Histogram.builder()
.name("order_processing_duration_seconds")
.help("Time to process one order")
.unit(Unit.SECONDS)
.register();
public static void main(String[] args) throws Exception {
JvmMetrics.builder().register(); // heap, GC, threads, classes, process
HTTPServer server = HTTPServer.builder().port(9400).buildAndStart();
// ... start the application
}
static void handle(Order o) {
long start = System.nanoTime();
try {
process(o);
ORDERS.labelValues("ok").inc();
} catch (RejectedException e) {
ORDERS.labelValues("rejected").inc();
} catch (RuntimeException e) {
ORDERS.labelValues("failed").inc();
throw e;
} finally {
LATENCY.observe((System.nanoTime() - start) / 1e9);
}
}
}Three habits are visible here. Metrics are static and created once; building a metric per request re-registers it and fails. Durations are recorded in seconds, the Prometheus base unit, never in milliseconds. And the label has a closed set of values decided by the code, not by the input. If the hot path calls labelValues() heavily, keep the returned child in a field instead of looking it up each time.
Spring Boot and Micrometer
In Spring Boot, add micrometer-registry-prometheus and the Actuator starter, then expose the endpoint. Boot registers JVM, process, HTTP server, connection pool and many other metrics automatically. Micrometer 1.13, which shipped with Spring Boot 3.3, moved its Prometheus registry onto client_java 1.x; older applications use the legacy simpleclient underneath.
# application.properties
management.endpoints.web.exposure.include=health,prometheus
management.metrics.tags.application=order-service
# publish histogram buckets for HTTP server latency so PromQL can compute quantiles
management.metrics.distribution.percentiles-histogram.http.server.requests=true@Component
class OrderHandler {
private final Timer latency;
private final MeterRegistry registry;
OrderHandler(MeterRegistry registry) {
this.registry = registry;
this.latency = Timer.builder("order.processing")
.publishPercentileHistogram()
.register(registry);
}
void handle(Order o) {
latency.record(() -> process(o));
registry.counter("orders.processed", "outcome", "ok").increment();
}
}Micrometer names use dots and are converted at export time: order.processing becomes order_processing_seconds with _count, _sum, _max and, because of the histogram flag, _bucket series. Micrometer's JVM metrics also have their own names, for example jvm_gc_pause_seconds, which differ from client_java's. Pick one source of JVM metrics per service and write dashboards against the names it actually exposes; scrape /actuator/prometheus once by hand and read it.
The JMX exporter for code you cannot change
For a JVM you cannot recompile, the JMX exporter runs as a Java agent inside the process, reads MBeans, and serves them on its own port. A YAML file of regex rules maps MBean names and attributes to metric names and labels.
# start the JVM with
# -javaagent:/opt/jmx_prometheus_javaagent.jar=9404:/etc/jmx/config.yaml
lowercaseOutputName: true
rules:
- pattern: 'kafka.server<type=BrokerTopicMetrics, name=(MessagesIn|BytesIn)PerSec><>Count'
name: kafka_server_brokertopicmetrics_$1_total
type: COUNTER
# no catch-all rule: unmatched MBeans are dropped instead of exportedPrefer the agent to the standalone exporter that connects over remote JMX: the agent sees process and JVM metrics directly and avoids an extra network hop. Write rules that whitelist what you need. A default catch-all rule on a Kafka broker or Cassandra node can export tens of thousands of series, many of them per topic or per table.
Choosing the metric type
| Type | Use for | Query with | Do not use for |
|---|---|---|---|
| Counter | Events and totals: requests, errors, bytes | rate(x_total[5m]) | Anything that can go down |
| Gauge | Current state: queue depth, pool in use, temperature | The value, max_over_time() | Event counts; you will miss events between scrapes |
| Histogram | Distributions: latency, payload size | histogram_quantile() over summed buckets | Exact percentiles of a single instance |
| Summary | Client-side quantiles from one instance | The quantile series directly | Anything you need to aggregate across instances |
Histograms are the right default for latency because buckets from many instances can be summed and a quantile computed over the whole fleet; quantiles from summaries cannot be averaged in any meaningful way. Classic histograms need bucket boundaries chosen up front, and each boundary is a series. Native histograms use exponential buckets chosen automatically, store the whole distribution as one series, and give better precision. client_java 1.x histograms maintain both representations by default; the native one is only transferred over the protobuf exposition format. On the server, native histograms are stable from Prometheus 3.8 but still opt-in through the scrape_native_histograms scrape setting, so check your server version before relying on them.
Labels and cardinality in Java code
Cardinality is the number of series a metric produces: the product of the distinct values of each label. A histogram with 12 buckets, 20 endpoints, 5 status codes and 3 methods is already about 3,600 series per instance before counting _sum and _count. Multiply by 50 pods and you have 180,000 series from one metric.
- Never put user ids, order ids, session ids, email addresses, full URLs or exception messages in labels. Put them in logs or traces, and link to them with exemplars if you need the connection.
- Use the route template,
/orders/{id}, not the raw path. Spring and most frameworks provide it; if yours does not, normalise in a filter and map unknown paths to one value. - Collapse status codes to classes such as
2xxand5xxunless you really alert on individual codes. - Do not use exception class names as labels without a fallback bucket; a new library version can introduce dozens.
Micrometer can enforce limits centrally with a MeterFilter, for example MeterFilter.maximumAllowableTags to cap the distinct values of a tag. The cardinality management guide covers finding and fixing an explosion once it has reached the server.
JVM metrics and the PromQL that reads them
JVM metrics are useful in a small set. The names below are the ones client_java's JvmMetrics exposes; adjust them if Micrometer is your source.
# Heap in use as a fraction of max (alert when sustained above about 0.9)
sum by (instance) (jvm_memory_used_bytes{area="heap"})
/ sum by (instance) (jvm_memory_max_bytes{area="heap"})
# Fraction of wall-clock time spent in GC over 5 minutes
sum by (instance) (rate(jvm_gc_collection_seconds_sum[5m]))
# Thread count growth, a leak signal
deriv(jvm_threads_current[30m])
# CPU used by the process, in cores
rate(process_cpu_seconds_total[5m])
# p99 request latency across the fleet from a classic histogram
histogram_quantile(0.99,
sum by (le) (rate(order_processing_duration_seconds_bucket[5m])))Two notes. jvm_gc_collection_seconds is a summary per collector, so its rate of _sum is seconds of GC per second; above 0.1 means the application loses more than a tenth of its time to collection. That reading holds for stop-the-world collectors such as G1, Parallel and Serial; ZGC and Shenandoah also report concurrent cycle time, so filter on the gc label to the pause collector before treating the sum as time lost. And a heap ratio of 0.9 is not by itself a problem for G1, which uses the heap it is given; what matters is whether it stays high after collections and whether GC time rises with it. The G1 collector article explains how to read those patterns.
Worked example: a latency regression caused by GC
A checkout service with 12 pods began breaching its p99 latency objective of 300 ms several times an hour. Average latency barely moved. The numbers below are illustrative, but the sequence is how such incidents are usually solved.
- The fleet-wide p99 from the histogram rose from 180 ms to 650 ms in short spikes. Splitting the same query
by (instance)showed the spikes moved between pods rather than hitting all of them at once, which rules out a shared downstream dependency. - On the affected pods,
rate(jvm_gc_collection_seconds_sum[1m])spiked to about 0.4 at the same moments, while request rate and CPU stayed flat. - The heap ratio after each collection had crept from 0.45 to 0.8 since a release two days earlier, and the old generation pool series showed steady growth: a cache added in that release had no size bound.
- The fix bounded the cache. The team then added an alert on GC time above 0.15 for ten minutes and a recording rule for p99 per instance, so the next time a single pod misbehaves the dashboard shows it immediately.
The lesson is not that GC causes latency. It is that per-instance breakdowns, a GC time rate and a post-collection heap trend together answer in minutes a question that averages hide entirely.
Failure modes and trade-offs
| Failure | Symptom | Fix |
|---|---|---|
| Unbounded labels | Server memory grows, scrapes slow, queries time out | Remove the label, normalise values, add a filter cap |
| Metric created per call | Registration errors or a leak of meter objects | Create metrics once as fields or static constants |
| Milliseconds in base metrics | Dashboards off by 1000x, mismatched buckets | Record seconds and bytes; convert only in display |
| Averaging summary quantiles | Fleet p99 looks healthy while one pod suffers | Use histograms and aggregate buckets |
| Slow scrape endpoint | Gaps and up == 0 under load | Keep collectors cheap; never do I/O inside a callback |
| Short-lived jobs | Batch finishes between scrapes; nothing recorded | Use a Pushgateway only for job completion metrics, or a longer-lived process |
| Duplicate JVM metrics | Two registries, conflicting names | One source of JVM metrics per process |
The main trade-off is resolution against cost. Shorter scrape intervals and more buckets show more detail but multiply storage. For most services a 15 to 30 second interval, a dozen buckets or a native histogram, and fewer than a few thousand series per instance is a comfortable budget. Prometheus metrics also tell you that something is wrong, not which request; pair them with tracing, for example through OpenTelemetry for Java, when you need per-request detail.
What to do next
- Scrape your service's endpoint by hand with curl and list every metric name and label; delete what nobody uses.
- Count series per instance with
count by (__name__) ({instance="host:port"})and investigate any metric above a few hundred. - Make sure latency is a histogram recorded in seconds, with route templates, not raw paths, as labels.
- Add the four JVM queries above to a dashboard: heap after GC, GC time rate, thread count trend and process CPU.
- Create alerts on symptoms users feel, error rate and p99 latency, and use JVM metrics for diagnosis rather than paging.
- Decide whether your server will scrape native histograms, and if so enable
scrape_native_histogramsfor the job. - For third-party JVMs, replace any catch-all JMX exporter rule with an explicit whitelist.