Cloud Monitoring is Google Cloud's metrics, dashboards and alerting service. Every Google Cloud product writes metrics into it automatically, and your own applications, VMs and Kubernetes workloads can write into the same store. Teams usually meet it through the console: a chart here, an alert there. That works until you need alerts that do not flap, dashboards that are correct across 40 projects, or a bill that does not surprise you. At that point you need the model underneath.
This article explains that model from first principles: what a time series is in Cloud Monitoring, how points are aligned and reduced, how metrics get in, how to query them with PromQL, how alerting policies and SLOs are built as code, and where the cost comes from. It ends with a worked example for a service on GKE and a checklist. Logs and traces have their own pages on Cloud Logging and Cloud Trace; here the subject is metrics.
The data model: descriptors, resources and series
A metric descriptor defines a metric type, such as compute.googleapis.com/instance/cpu/utilization, along with its kind, value type, unit and allowed labels. Google-defined types are prefixed by the service domain; your own use prefixes such as custom.googleapis.com/, workload.googleapis.com/ for Ops Agent and OpenTelemetry metrics, or prometheus.googleapis.com/ for Managed Prometheus.
A monitored resource is the thing being measured, with its own type and labels: gce_instance with instance_id and zone, k8s_container with cluster, namespace, pod and container names, or generic_task for anything else. A time series is identified by the metric type, the values of its metric labels, the resource type and the values of the resource labels. Change any label value and you have created a new series. That rule is the source of both the model's power and most of its cost.
The metric kind says how to read the points. A GAUGE is a value at an instant, such as memory in use. A DELTA is the change over an interval, such as requests in the last minute. A CUMULATIVE is a running total since a start time, such as bytes sent since the process began, and the store handles resets. The value type can be boolean, integer, double, string or distribution. Distributions carry a count, mean, sum of squared deviations and bucket counts, which is how latency histograms are stored, and why you can compute a 99th percentile across 500 pods without shipping raw samples.
The architecture and metrics scopes
Every series belongs to a Google Cloud project, but you query through a metrics scope: the set of projects whose series a given scoping project can read. Create a dedicated monitoring project per environment or per team, add the service projects to its scope, and keep dashboards, alerting policies, uptime checks and notification channels there. This separates who can change alerts from who can deploy workloads, and gives one place to see a system that spans projects. Keep scopes aligned with ownership. A scope that contains everything becomes a scope where every alert policy fires for resources nobody on the team owns.
Getting metrics in
System metrics from Google Cloud services, such as Compute Engine, Cloud SQL, load balancers, Pub/Sub and GKE, arrive automatically and are free. Check what exists before instrumenting anything: load balancer request counts and latency distributions often answer the question without application code. The Ops Agent on VMs adds host metrics such as memory and disk usage that the hypervisor cannot see, and scrapes common applications. On GKE, Google Cloud Managed Service for Prometheus runs collectors that you configure with PodMonitoring resources; your pods keep exposing /metrics exactly as for self-run Prometheus, and the data lands in Cloud Monitoring, queryable with PromQL. For how Prometheus itself stores and queries data, see the Prometheus deep dive. Log-based metrics turn matching log entries into counters or distributions, useful for systems you cannot instrument. And the API, directly or through an OpenTelemetry exporter, writes custom metrics.
import time
from google.cloud import monitoring_v3
client = monitoring_v3.MetricServiceClient()
project = "projects/shop-prod"
series = monitoring_v3.TimeSeries()
series.metric.type = "custom.googleapis.com/checkout/queue_depth"
series.metric.labels["queue"] = "payments"
series.resource.type = "generic_task"
series.resource.labels.update({"project_id": "shop-prod", "location": "europe-west1",
"namespace": "checkout", "job": "worker", "task_id": "7"})
now = time.time()
point = monitoring_v3.Point({
"interval": {"end_time": {"seconds": int(now), "nanos": int((now % 1) * 1e9)}},
"value": {"int64_value": 42},
})
series.points = [point]
client.create_time_series(name=project, time_series=[series]) # up to 200 series per callThe limits worth knowing when designing custom metrics: at most one point every 5 seconds per series, at most 30 labels per custom metric descriptor, at most 200 series per create call, and 10,000 custom metric descriptors per project. Custom, agent and external metrics are retained for 24 months. In practice, aggregate in-process and write every 60 seconds; writing every 5 seconds multiplies ingestion for resolution that most alerts cannot use.
Querying: align first, then reduce
Every query in Cloud Monitoring, whether built in the console, written as a filter in an API call or expressed in PromQL, does two steps. Alignment turns each series' irregular points into one value per alignment period, which must be at least 60 seconds for alerting aggregations, using a per-series aligner such as mean, max, rate, delta or a percentile. Reduction then combines aligned series across labels, such as a sum across all pods grouped by region. Reversing the order in your head causes most wrong charts. For example, averaging per-pod p99 latencies gives a meaningless number; the correct approach is to sum the distributions and then take the percentile.
PromQL is the recommended query language. Metric types are mapped to PromQL names by replacing the first / with : and every other special character with _, so CPU utilisation becomes compute_googleapis_com:instance_cpu_utilization. When a metric can be written against several resource types, you must add a monitored_resource label matcher. The older Monitoring Query Language, MQL, is deprecated: console support ended on 22 July 2025. Existing MQL charts and policies keep working and the API still accepts them, but new work should use PromQL.
# 99th percentile checkout latency per region, from a Managed Prometheus histogram
histogram_quantile(0.99,
sum by (le, location) (rate(http_server_duration_seconds_bucket{job="checkout"}[5m])))
# Load balancer 5xx ratio, from a Google Cloud system metric
sum(rate(loadbalancing_googleapis_com:https_request_count{monitored_resource="https_lb_rule",
response_code_class="500"}[5m]))
/
sum(rate(loadbalancing_googleapis_com:https_request_count{monitored_resource="https_lb_rule"}[5m]))
Alerting policies as code
An alerting policy has conditions, a combiner (AND or OR across conditions), notification channels, documentation and an alert strategy. Condition types include a metric threshold (conditionThreshold), metric absence (conditionAbsent) and PromQL (conditionPrometheusQueryLanguage). A threshold condition has a filter, aggregations, a comparison of COMPARISON_GT or COMPARISON_LT, a threshold value, a duration for which the series must violate it, and a trigger saying how many series must fail. A PromQL condition takes a query, an optional duration and an evaluation interval that must be a multiple of 30 seconds.
{
"displayName": "checkout: 5xx ratio above 2% for 5 minutes",
"combiner": "OR",
"conditions": [{
"displayName": "5xx ratio > 0.02",
"conditionPrometheusQueryLanguage": {
"query": "sum(rate(http_requests_total{job=\"checkout\",code=~\"5..\"}[5m])) / sum(rate(http_requests_total{job=\"checkout\"}[5m])) > 0.02",
"duration": "300s",
"evaluationInterval": "60s"
}
}],
"alertStrategy": {"autoClose": "1800s"},
"documentation": {
"content": "Runbook: go/checkout-5xx. First check the last deploy and the payments dependency dashboard.",
"mimeType": "text/markdown"
},
"notificationChannels": ["projects/shop-monitoring/notificationChannels/1234567890"]
}Store policies like this in Git and apply them with Terraform or gcloud, never by clicking; console-built alerts drift and nobody can review them. Always pair an important threshold with an absence condition, because a threshold on a series that stops reporting is silently never true. Use documentation to put the runbook link and first diagnostic step in every notification. Use autoClose so incidents on vanished resources do not stay open for weeks. The quotas are 2,000 alerting policies per metrics scope and six conditions per metric-based policy, and hitting either is a sign that alerts need consolidating, not a quota increase.
SLOs and burn-rate alerting
Threshold alerts on causes, such as CPU or queue depth, page people for things users never notice. Cloud Monitoring's service monitoring lets you define a service, attach SLOs to it and alert on how fast the error budget is burning. A request-based SLI compares a good-request filter against a total-request filter.
{
"displayName": "checkout availability 99.9% / 28d",
"serviceLevelIndicator": {
"requestBased": {
"goodTotalRatio": {
"totalServiceFilter": "resource.type=https_lb_rule metric.type=\"loadbalancing.googleapis.com/https/request_count\"",
"goodServiceFilter": "resource.type=https_lb_rule metric.type=\"loadbalancing.googleapis.com/https/request_count\" metric.label.response_code_class=200"
}
}
},
"goal": 0.999,
"rollingPeriod": "2419200s"
}A burn-rate alert is then an ordinary threshold condition whose filter is select_slo_burn_rate("projects/PROJECT/services/SERVICE/serviceLevelObjectives/SLO", "60m"). A burn rate of 1 spends the budget exactly over the SLO period. Google's documentation suggests starting with a fast-burn policy at 10 times the baseline over a 1 or 2 hour lookback, which pages, and a slow-burn policy at 2 times over 24 hours, which opens a ticket. The reasoning behind multi-window burn rates is covered in burn-rate alerting. The good filter here counts only 2xx responses, which also treats 3xx and 4xx as bad. Decide deliberately which classes count against the budget; client errors usually should not.
Cardinality and cost
System metrics from Google Cloud services are free. Custom, agent, external and Managed Prometheus metrics are chargeable, some billed by bytes ingested and Managed Prometheus by samples ingested, and alerting and API reads have their own pricing; check the current pricing page rather than any figure copied into a design doc. Whatever the unit, cost scales with series count multiplied by write frequency.
Series count is the product of label cardinalities. A request counter with labels for 30 endpoints, 5 status classes, 3 methods and 400 pods is 180,000 series; add a customer_id label with 20,000 values and you have created a metric that is unaffordable to store and too slow to query. Rules: never put unbounded identifiers such as user, order or trace ids in labels; use Managed Prometheus relabelling to drop unused metrics at the collector; keep histogram buckets few; and sample every 60 seconds unless an alert needs more. Review the metrics management page monthly for the top metrics by volume. The same discipline applies to GKE clusters, where one badly labelled exporter on every node multiplies quickly.
Worked example: monitoring a checkout service on GKE
A checkout service runs on GKE in two regions behind a global external Application Load Balancer. Step 1: the team adds the monitoring project shop-monitoring and puts the two regional projects in its metrics scope. Step 2: the pods already expose Prometheus histograms, so a PodMonitoring resource scrapes them every 60 seconds. Relabelling drops dozens of runtime metrics nobody queries, and the latency histogram goes from 30 buckets to 12, which cuts series by more than half.
Step 3: an SLO is defined on the load balancer request count, 99.9 percent over 28 days, because the load balancer sees failures the pods never log, such as connection resets. Step 4: two burn-rate policies are created in Terraform, fast at 10 over 60 minutes routed to PagerDuty, slow at 2 over 24 hours routed to a ticket queue, with runbook links in documentation. Step 5: absence conditions cover the scrape itself, so a broken collector pages once rather than silently hiding everything.
Two weeks later a bad deploy in one region makes 20 percent of that region's requests fail, about 10 percent of all traffic. Against a 0.1 percent budget that is a burn rate of about 100, so the 60-minute average passes 10 after roughly six minutes and the fast-burn alert fires with the runbook link. By the time of the rollback at minute 15, about 4 percent of the 28-day budget is gone. A milder regression at a burn rate of 15 would take about 40 minutes to trip the same policy, which is why many teams add a second, shorter window. The regional breakdown on the linked dashboard shows the problem is in one region, the deploy is rolled back, and the slow-burn policy never fires. Before the SLOs existed, the same incident produced 30 CPU and pod-restart alerts and no clear signal.
Failure modes
- Silent absence. A threshold on a series that stops reporting never fires. Add absence conditions for critical signals and for the collectors themselves.
- Wrong aligner. Using mean on a CUMULATIVE counter, or averaging percentiles, gives charts that look plausible and are wrong. Use rate or delta for counters and aggregate distributions before taking percentiles.
- Flapping. A zero duration on a noisy metric pages on every spike. Use a duration, or a burn rate over a window.
- Cardinality explosion. A new label deployed on Friday doubles the bill by Monday. Review label changes like schema changes.
- Console drift. Hand-edited policies diverge from Git and lose their runbooks. Apply policies from code only.
- Scope confusion. Policies in a service project instead of the monitoring project, or scopes that include unrelated projects, route alerts to the wrong team.
What to do next
- Create a monitoring project per environment and add the right service projects to its metrics scope.
- Inventory the free system metrics for your services, especially load balancer latency and request counts, before writing custom metrics.
- Instrument with Prometheus client libraries or OpenTelemetry, scrape with Managed Service for Prometheus on GKE, and write custom metrics every 60 seconds with bounded labels.
- Define one request-based SLO per user-facing service and add fast and slow burn-rate policies.
- Move every alerting policy into Terraform with runbook documentation, autoClose and a matching absence condition.
- Migrate any remaining MQL charts and policies to PromQL, and review top metrics by volume on the metrics management page monthly.