Almost every AWS workload sends data to Amazon CloudWatch whether or not anyone looks at it. EC2, Lambda, RDS and load balancers publish metrics automatically, Lambda writes its logs there, and the first alarm most teams create is a CloudWatch alarm. Yet most teams use only a fraction of it well: alarms that flap or never fire, log groups that keep everything forever, and a custom-metrics bill that grows faster than traffic.

This page explains how CloudWatch actually models and evaluates data, so that you can design metrics, alarms and logs on purpose. It covers the metric data model and retention, the two ways to publish custom metrics, exactly how an alarm decides to change state, logs and log classes, Logs Insights, and a worked example for a checkout service. It ends with failure modes and a checklist you can apply this week.

Advertisement

The data model: what a metric really is

A CloudWatch metric is a time series identified by three things: a namespace such as AWS/EC2 or Shop/Checkout, a metric name, and a set of dimensions, which are name-value pairs such as Service=checkout. Change any one of them and you have a different metric. A metric can have up to 30 dimensions, and the identity is the exact combination.

Exact matching has two big consequences. First, CloudWatch does not add up across dimensions for you when you read a single metric: if you publish latency with Service and Route, there is no automatic metric for Service alone. You either publish that combination too, or you aggregate at query time with Metrics Insights, whose SQL-like queries such as SELECT AVG(CPUUtilization) FROM SCHEMA("AWS/EC2", InstanceId) GROUP BY InstanceId can group across metrics. Second, an alarm or graph that names the wrong dimension set finds no data at all, rather than an error.

Each metric holds datapoints with a timestamp, a value and a unit. When you read it, you ask for a statistic over a period: Sum, Average, Minimum, Maximum, SampleCount, or a percentile such as p99. Percentiles matter because Average hides the slow tail that users actually feel.

CloudWatch: metrics and logs flow in, alarms decide, actions fireAWS servicesEC2, Lambda, RDS, ALBYour codePutMetricDataYour codeEMF lines to logsCloudWatch agenthost metrics, log filesMetrics storenamespace, name, dimensionsLogslog groups, streams, classEMF and metric filtersAlarmsperiod, M of N, missing dataComposite alarmsALARM(a) AND ALARM(b)Logs Insightsquery by GB scannedSubscriptions, exportLambda, Firehose, S3Alarm actions: SNS to people or tools, Auto Scaling policies, EC2 actions. Dashboards read the same metrics.Every distinct namespace, name and dimension set is a separate billed metric, so labels decide the bill.
Data enters as metrics or logs. Logs can produce metrics through EMF and metric filters. Alarms evaluate metrics and trigger actions; Logs Insights queries the raw events.

Resolution and retention

Standard-resolution metrics have one-minute granularity. High-resolution custom metrics, published with StorageResolution=1, keep one-second granularity and allow sub-minute alarm periods at a higher price. CloudWatch then rolls data up as it ages:

PeriodKept for
Under 60 seconds (high resolution)3 hours
60 seconds15 days
5 minutes63 days
1 hour455 days (15 months)

After 15 days you can no longer see per-minute detail for an incident, and after 63 days you only have hourly data. If you need longer or finer history for capacity planning, export it: metric streams can send metrics continuously to Amazon Data Firehose, and from there to S3 or another monitoring tool.

Advertisement

Publishing custom metrics: API or embedded format

There are two ways for your code to create metrics. The first is the PutMetricData API. One request can carry up to 1,000 metrics and up to 1 MB, and it can be gzip-compressed. Batch observations with Values and Counts instead of making a call per request:

import boto3

cw = boto3.client("cloudwatch")

# Up to 1,000 metrics and 1 MB per request. Batch, do not call once per event.
cw.put_metric_data(
    Namespace="Shop/Checkout",
    MetricData=[{
        "MetricName": "PaymentLatency",
        "Dimensions": [{"Name": "Service", "Value": "checkout"},
                       {"Name": "Route", "Value": "/pay"}],
        "Values": [182, 240, 95, 1210],     # many observations in one datum
        "Counts": [1, 1, 3, 1],
        "Unit": "Milliseconds",
        "StorageResolution": 60,            # 1 = high resolution, billed and retained differently
    }],
)

The second is the Embedded Metric Format (EMF). Your code writes one structured JSON log line, and CloudWatch Logs extracts metrics from it as the line is ingested. In Lambda you simply print it. On EC2 or containers, the CloudWatch agent or a compatible collector forwards it. The _aws block declares which fields are metrics and which are dimensions, and every other field stays in the log event for searching:

{"_aws": {"Timestamp": 1790913600000,
          "CloudWatchMetrics": [{"Namespace": "Shop/Checkout",
                                 "Dimensions": [["Service", "Route"]],
                                 "Metrics": [{"Name": "PaymentLatency", "Unit": "Milliseconds"},
                                             {"Name": "PaymentErrors", "Unit": "Count"}]}]},
 "Service": "checkout", "Route": "/pay",
 "PaymentLatency": 182, "PaymentErrors": 0,
 "orderId": "o-88412", "customerTier": "gold", "traceId": "1-66fc-ab12"}

This line creates two metrics with the dimension set Service, Route, and the high-cardinality fields orderId and traceId stay in the log where they cost nothing extra as metrics. That split is the main design rule for custom metrics: dimensions are for values with a small fixed set of options, such as service, route or region. Customer ids, order ids and request ids belong in logs. Put one of those in a dimension and every distinct value becomes its own billed metric.

EMF suits applications that already log every request, because metrics and context arrive together without extra API calls on the request path. The API suits aggregated values you compute yourself, such as a queue depth sampled every minute.

How an alarm decides

An alarm watches one metric, or a metric math expression, and has three states: OK, ALARM and INSUFFICIENT_DATA. Each evaluation computes the chosen statistic for each period and compares it with the threshold. Three settings decide what counts as a breach:

  • Period: the length of each datapoint, for example 60 seconds.
  • EvaluationPeriods (N): how many of the most recent periods to look at.
  • DatapointsToAlarm (M): how many of those N must breach. M of N alarms ignore a single spike but still catch a sustained problem quickly.

The fourth setting, TreatMissingData, decides what an empty period means. missing (the default) does not count it either way, notBreaching treats it as good, breaching treats it as bad, and ignore keeps the current state. Choose it per metric. For a latency metric that only exists when there is traffic, missing data at 3 a.m. is fine, so use notBreaching. For a heartbeat metric whose silence means the job died, use breaching, because missing is the failure.

Alarm actions can publish to an SNS topic, trigger an Auto Scaling policy, or stop, terminate, reboot or recover an EC2 instance. Composite alarms combine other alarms with a rule language using ALARM(), OK(), AND, OR and NOT, which is the main tool against alert fatigue: page only when the user-facing symptom and a cause agree, and send the parts to a ticket queue. For metrics with daily cycles, anomaly detection alarms compare against a learned band instead of a fixed threshold.

Logs: groups, retention and classes

Logs are organised as log groups (one per application or function) containing log streams (one per instance, container or Lambda execution environment). Two settings belong on every log group from the day it is created. The first is retention: a new log group keeps data forever unless you set a retention period, and storage is billed for as long as it is kept. The second is the log class, which cannot be changed after the group exists.

FeatureStandardInfrequent Access
Ingestion priceStandardLower per GB (storage and Insights prices are the same)
Logs Insights queriesYesYes, most commands
Metric filters and Embedded Metric FormatYesNo
Subscription filtersYesNo
Live Tail, GetLogEvents, FilterLogEventsYesNo; read with Logs Insights
Anomaly detection, field indexesYesNo
Export to S3, KMS encryption, cross-accountYesYes

The table is the rule: any log group that produces metrics through EMF or metric filters, or that streams to another system, must be Standard. Infrequent Access is for audit and debug logs you search rarely and after the fact.

Metric filters turn log patterns into metrics without code changes, for example counting lines that match { $.level = "ERROR" }. Subscription filters stream matching events to Lambda, Kinesis Data Streams, Firehose or OpenSearch in near real time, which is how logs reach a central security account or a search cluster. See Kinesis architecture if you are building that pipeline.

Logs Insights

Logs Insights is a query language over log groups that discovers JSON fields automatically. Queries are piped stages:

fields @timestamp, Route, PaymentLatency, orderId
| filter Service = "checkout" and PaymentLatency > 800
| stats count(*) as slow, pct(PaymentLatency, 99) as p99 by Route, bin(5m)
| sort slow desc
| limit 20

You pay by the volume of data scanned, so narrow the time range and the set of log groups before anything else. Field indexes on Standard log groups can reduce what a query scans when you filter on an indexed field. Insights is ideal for the question after the alarm (which route, which customers, since when), not for dashboards refreshed every minute. Those should read metrics.

Worked example: alerting a checkout service

A checkout service runs on AWS Lambda. The goal is to page a human only when customers are actually affected, and to leave enough data to diagnose it. The design: every request writes one EMF line like the one above, in a Standard log group with 30-day retention. That yields PaymentLatency and PaymentErrors per route at a cost of a handful of metrics, while order ids stay searchable in the log.

cw.put_metric_alarm(
    AlarmName="checkout-pay-p99-high",
    Namespace="Shop/Checkout",
    MetricName="PaymentLatency",
    Dimensions=[{"Name": "Service", "Value": "checkout"},
                {"Name": "Route", "Value": "/pay"}],      # must match the published set exactly
    ExtendedStatistic="p99",
    Period=60,
    EvaluationPeriods=5,
    DatapointsToAlarm=3,                                  # 3 bad minutes out of the last 5
    Threshold=800,
    ComparisonOperator="GreaterThanThreshold",
    TreatMissingData="notBreaching",                      # no traffic is not an outage here
    AlarmActions=[TICKET_TOPIC_ARN],
    OKActions=[TICKET_TOPIC_ARN],
)

cw.put_composite_alarm(
    AlarmName="checkout-pay-degraded",
    AlarmRule='ALARM("checkout-pay-p99-high") AND ALARM("checkout-pay-error-rate-high")',
    AlarmActions=[PAGER_TOPIC_ARN],                       # page only when both are true
)

Walk through an incident. At 14:02 the payment provider slows down and p99 on /pay goes from 300 ms to 1.4 s. The minute datapoints for 14:02, 14:03 and 14:04 breach, so at the evaluation after 14:04 the p99 alarm has three breaching datapoints out of five and moves to ALARM, opening a ticket. Errors are still low, so the composite stays OK and nobody is woken. At 14:06 timeouts start, the error-rate alarm (an error ratio built with metric math, 100 * errors / requests) also enters ALARM, and the composite pages through SNS. The on-call engineer runs the Insights query, sees the slowness is limited to one route and started at 14:02, and switches the provider. One spike at 14:30 breaches a single minute and does nothing, because one breach in five is below the M of N rule.

Failure modes

  • Alarm stuck in INSUFFICIENT_DATA. The alarm's dimensions do not match the published set exactly, or the period is shorter than how often data arrives. Copy dimensions from the metric in the console, and use a period at least as long as the publishing interval.
  • Alarms on Average. The average stays healthy while one user in a hundred waits ten seconds. Alarm on p99 or p95 for latency.
  • Flapping. A threshold close to normal values with 1 of 1 evaluation. Use M of N and a threshold based on observed behaviour.
  • Cardinality explosion. A request id, user id or pod name used as a dimension creates thousands of metrics. Keep those in logs.
  • Logs kept forever. Groups created implicitly, such as by a new Lambda function, have no retention. Enforce retention with infrastructure as code or an EventBridge rule that sets it on creation.
  • EMF metrics that never appear. The log group is Infrequent Access, which does not extract EMF, or the JSON is malformed, or the dimension field is missing from the line.
  • Silent heartbeats. A batch job's success metric uses notBreaching, so when the job stops running the alarm stays OK. Missing data must be breaching for that kind of metric.

Trade-offs

ChoiceGainCost
EMF instead of PutMetricDataNo API calls on the request path; metrics and context togetherIngestion cost for every line; requires Standard class
High-resolution metricsSecond-level detail and faster alarmsHigher price and 3-hour retention at full detail
Composite alarms for pagingFewer, more meaningful pagesMore alarms to maintain
Infrequent Access logsLower ingestion costNo metric or subscription filters, no Live Tail
CloudWatch versus Prometheus or a SaaS toolNo servers, native AWS metrics and actionsPer-metric pricing punishes high cardinality; weaker query language

Many teams keep CloudWatch for AWS service metrics, alarms that trigger AWS actions and Lambda logs, and send high-cardinality application telemetry elsewhere. That is a reasonable split, as long as the paging alarms live in one place. Scaling on CloudWatch alarms is covered in Auto Scaling Groups.

What to do next

  1. List your log groups with no retention and set a retention period on each one.
  2. Check every log group that uses EMF, metric filters or subscriptions is in the Standard class.
  3. Audit custom metric dimensions and move any per-user or per-request value into logs.
  4. Change latency alarms from Average to p99 and use M of N evaluation.
  5. Set TreatMissingData deliberately on every alarm, with breaching for heartbeats.
  6. Create one composite alarm per user-facing service for paging, and route the rest to tickets.
  7. Save the Logs Insights queries your on-call engineers need, with narrow time ranges.
  8. If you need history beyond 15 months or finer than hourly after 63 days, set up a metric stream.
Key takeaway: A CloudWatch metric is an exact namespace, name and dimension set, rolled up over time from one-minute detail to hourly data kept for 15 months. Publish with EMF or batched PutMetricData, keep high-cardinality values in logs, and alarm on percentiles with M of N evaluation and a deliberate missing-data policy. Give every log group a retention period and the right class, page through composite alarms, and use Logs Insights to explain the alarm.