Almost every AWS workload sends data to Amazon CloudWatch whether or not anyone looks at it. EC2, Lambda, RDS and load balancers publish metrics automatically, Lambda writes its logs there, and the first alarm most teams create is a CloudWatch alarm. Yet most teams use only a fraction of it well: alarms that flap or never fire, log groups that keep everything forever, and a custom-metrics bill that grows faster than traffic.
This page explains how CloudWatch actually models and evaluates data, so that you can design metrics, alarms and logs on purpose. It covers the metric data model and retention, the two ways to publish custom metrics, exactly how an alarm decides to change state, logs and log classes, Logs Insights, and a worked example for a checkout service. It ends with failure modes and a checklist you can apply this week.
The data model: what a metric really is
A CloudWatch metric is a time series identified by three things: a namespace such as AWS/EC2 or Shop/Checkout, a metric name, and a set of dimensions, which are name-value pairs such as Service=checkout. Change any one of them and you have a different metric. A metric can have up to 30 dimensions, and the identity is the exact combination.
Exact matching has two big consequences. First, CloudWatch does not add up across dimensions for you when you read a single metric: if you publish latency with Service and Route, there is no automatic metric for Service alone. You either publish that combination too, or you aggregate at query time with Metrics Insights, whose SQL-like queries such as SELECT AVG(CPUUtilization) FROM SCHEMA("AWS/EC2", InstanceId) GROUP BY InstanceId can group across metrics. Second, an alarm or graph that names the wrong dimension set finds no data at all, rather than an error.
Each metric holds datapoints with a timestamp, a value and a unit. When you read it, you ask for a statistic over a period: Sum, Average, Minimum, Maximum, SampleCount, or a percentile such as p99. Percentiles matter because Average hides the slow tail that users actually feel.
Resolution and retention
Standard-resolution metrics have one-minute granularity. High-resolution custom metrics, published with StorageResolution=1, keep one-second granularity and allow sub-minute alarm periods at a higher price. CloudWatch then rolls data up as it ages:
| Period | Kept for |
|---|---|
| Under 60 seconds (high resolution) | 3 hours |
| 60 seconds | 15 days |
| 5 minutes | 63 days |
| 1 hour | 455 days (15 months) |
After 15 days you can no longer see per-minute detail for an incident, and after 63 days you only have hourly data. If you need longer or finer history for capacity planning, export it: metric streams can send metrics continuously to Amazon Data Firehose, and from there to S3 or another monitoring tool.
Publishing custom metrics: API or embedded format
There are two ways for your code to create metrics. The first is the PutMetricData API. One request can carry up to 1,000 metrics and up to 1 MB, and it can be gzip-compressed. Batch observations with Values and Counts instead of making a call per request:
import boto3
cw = boto3.client("cloudwatch")
# Up to 1,000 metrics and 1 MB per request. Batch, do not call once per event.
cw.put_metric_data(
Namespace="Shop/Checkout",
MetricData=[{
"MetricName": "PaymentLatency",
"Dimensions": [{"Name": "Service", "Value": "checkout"},
{"Name": "Route", "Value": "/pay"}],
"Values": [182, 240, 95, 1210], # many observations in one datum
"Counts": [1, 1, 3, 1],
"Unit": "Milliseconds",
"StorageResolution": 60, # 1 = high resolution, billed and retained differently
}],
)The second is the Embedded Metric Format (EMF). Your code writes one structured JSON log line, and CloudWatch Logs extracts metrics from it as the line is ingested. In Lambda you simply print it. On EC2 or containers, the CloudWatch agent or a compatible collector forwards it. The _aws block declares which fields are metrics and which are dimensions, and every other field stays in the log event for searching:
{"_aws": {"Timestamp": 1790913600000,
"CloudWatchMetrics": [{"Namespace": "Shop/Checkout",
"Dimensions": [["Service", "Route"]],
"Metrics": [{"Name": "PaymentLatency", "Unit": "Milliseconds"},
{"Name": "PaymentErrors", "Unit": "Count"}]}]},
"Service": "checkout", "Route": "/pay",
"PaymentLatency": 182, "PaymentErrors": 0,
"orderId": "o-88412", "customerTier": "gold", "traceId": "1-66fc-ab12"}This line creates two metrics with the dimension set Service, Route, and the high-cardinality fields orderId and traceId stay in the log where they cost nothing extra as metrics. That split is the main design rule for custom metrics: dimensions are for values with a small fixed set of options, such as service, route or region. Customer ids, order ids and request ids belong in logs. Put one of those in a dimension and every distinct value becomes its own billed metric.
EMF suits applications that already log every request, because metrics and context arrive together without extra API calls on the request path. The API suits aggregated values you compute yourself, such as a queue depth sampled every minute.
How an alarm decides
An alarm watches one metric, or a metric math expression, and has three states: OK, ALARM and INSUFFICIENT_DATA. Each evaluation computes the chosen statistic for each period and compares it with the threshold. Three settings decide what counts as a breach:
- Period: the length of each datapoint, for example 60 seconds.
- EvaluationPeriods (N): how many of the most recent periods to look at.
- DatapointsToAlarm (M): how many of those N must breach. M of N alarms ignore a single spike but still catch a sustained problem quickly.
The fourth setting, TreatMissingData, decides what an empty period means. missing (the default) does not count it either way, notBreaching treats it as good, breaching treats it as bad, and ignore keeps the current state. Choose it per metric. For a latency metric that only exists when there is traffic, missing data at 3 a.m. is fine, so use notBreaching. For a heartbeat metric whose silence means the job died, use breaching, because missing is the failure.
Alarm actions can publish to an SNS topic, trigger an Auto Scaling policy, or stop, terminate, reboot or recover an EC2 instance. Composite alarms combine other alarms with a rule language using ALARM(), OK(), AND, OR and NOT, which is the main tool against alert fatigue: page only when the user-facing symptom and a cause agree, and send the parts to a ticket queue. For metrics with daily cycles, anomaly detection alarms compare against a learned band instead of a fixed threshold.
Logs: groups, retention and classes
Logs are organised as log groups (one per application or function) containing log streams (one per instance, container or Lambda execution environment). Two settings belong on every log group from the day it is created. The first is retention: a new log group keeps data forever unless you set a retention period, and storage is billed for as long as it is kept. The second is the log class, which cannot be changed after the group exists.
| Feature | Standard | Infrequent Access |
|---|---|---|
| Ingestion price | Standard | Lower per GB (storage and Insights prices are the same) |
| Logs Insights queries | Yes | Yes, most commands |
| Metric filters and Embedded Metric Format | Yes | No |
| Subscription filters | Yes | No |
| Live Tail, GetLogEvents, FilterLogEvents | Yes | No; read with Logs Insights |
| Anomaly detection, field indexes | Yes | No |
| Export to S3, KMS encryption, cross-account | Yes | Yes |
The table is the rule: any log group that produces metrics through EMF or metric filters, or that streams to another system, must be Standard. Infrequent Access is for audit and debug logs you search rarely and after the fact.
Metric filters turn log patterns into metrics without code changes, for example counting lines that match { $.level = "ERROR" }. Subscription filters stream matching events to Lambda, Kinesis Data Streams, Firehose or OpenSearch in near real time, which is how logs reach a central security account or a search cluster. See Kinesis architecture if you are building that pipeline.
Logs Insights
Logs Insights is a query language over log groups that discovers JSON fields automatically. Queries are piped stages:
fields @timestamp, Route, PaymentLatency, orderId
| filter Service = "checkout" and PaymentLatency > 800
| stats count(*) as slow, pct(PaymentLatency, 99) as p99 by Route, bin(5m)
| sort slow desc
| limit 20You pay by the volume of data scanned, so narrow the time range and the set of log groups before anything else. Field indexes on Standard log groups can reduce what a query scans when you filter on an indexed field. Insights is ideal for the question after the alarm (which route, which customers, since when), not for dashboards refreshed every minute. Those should read metrics.
Worked example: alerting a checkout service
A checkout service runs on AWS Lambda. The goal is to page a human only when customers are actually affected, and to leave enough data to diagnose it. The design: every request writes one EMF line like the one above, in a Standard log group with 30-day retention. That yields PaymentLatency and PaymentErrors per route at a cost of a handful of metrics, while order ids stay searchable in the log.
cw.put_metric_alarm(
AlarmName="checkout-pay-p99-high",
Namespace="Shop/Checkout",
MetricName="PaymentLatency",
Dimensions=[{"Name": "Service", "Value": "checkout"},
{"Name": "Route", "Value": "/pay"}], # must match the published set exactly
ExtendedStatistic="p99",
Period=60,
EvaluationPeriods=5,
DatapointsToAlarm=3, # 3 bad minutes out of the last 5
Threshold=800,
ComparisonOperator="GreaterThanThreshold",
TreatMissingData="notBreaching", # no traffic is not an outage here
AlarmActions=[TICKET_TOPIC_ARN],
OKActions=[TICKET_TOPIC_ARN],
)
cw.put_composite_alarm(
AlarmName="checkout-pay-degraded",
AlarmRule='ALARM("checkout-pay-p99-high") AND ALARM("checkout-pay-error-rate-high")',
AlarmActions=[PAGER_TOPIC_ARN], # page only when both are true
)Walk through an incident. At 14:02 the payment provider slows down and p99 on /pay goes from 300 ms to 1.4 s. The minute datapoints for 14:02, 14:03 and 14:04 breach, so at the evaluation after 14:04 the p99 alarm has three breaching datapoints out of five and moves to ALARM, opening a ticket. Errors are still low, so the composite stays OK and nobody is woken. At 14:06 timeouts start, the error-rate alarm (an error ratio built with metric math, 100 * errors / requests) also enters ALARM, and the composite pages through SNS. The on-call engineer runs the Insights query, sees the slowness is limited to one route and started at 14:02, and switches the provider. One spike at 14:30 breaches a single minute and does nothing, because one breach in five is below the M of N rule.
Failure modes
- Alarm stuck in INSUFFICIENT_DATA. The alarm's dimensions do not match the published set exactly, or the period is shorter than how often data arrives. Copy dimensions from the metric in the console, and use a period at least as long as the publishing interval.
- Alarms on Average. The average stays healthy while one user in a hundred waits ten seconds. Alarm on p99 or p95 for latency.
- Flapping. A threshold close to normal values with 1 of 1 evaluation. Use M of N and a threshold based on observed behaviour.
- Cardinality explosion. A request id, user id or pod name used as a dimension creates thousands of metrics. Keep those in logs.
- Logs kept forever. Groups created implicitly, such as by a new Lambda function, have no retention. Enforce retention with infrastructure as code or an EventBridge rule that sets it on creation.
- EMF metrics that never appear. The log group is Infrequent Access, which does not extract EMF, or the JSON is malformed, or the dimension field is missing from the line.
- Silent heartbeats. A batch job's success metric uses notBreaching, so when the job stops running the alarm stays OK. Missing data must be breaching for that kind of metric.
Trade-offs
| Choice | Gain | Cost |
|---|---|---|
| EMF instead of PutMetricData | No API calls on the request path; metrics and context together | Ingestion cost for every line; requires Standard class |
| High-resolution metrics | Second-level detail and faster alarms | Higher price and 3-hour retention at full detail |
| Composite alarms for paging | Fewer, more meaningful pages | More alarms to maintain |
| Infrequent Access logs | Lower ingestion cost | No metric or subscription filters, no Live Tail |
| CloudWatch versus Prometheus or a SaaS tool | No servers, native AWS metrics and actions | Per-metric pricing punishes high cardinality; weaker query language |
Many teams keep CloudWatch for AWS service metrics, alarms that trigger AWS actions and Lambda logs, and send high-cardinality application telemetry elsewhere. That is a reasonable split, as long as the paging alarms live in one place. Scaling on CloudWatch alarms is covered in Auto Scaling Groups.
What to do next
- List your log groups with no retention and set a retention period on each one.
- Check every log group that uses EMF, metric filters or subscriptions is in the Standard class.
- Audit custom metric dimensions and move any per-user or per-request value into logs.
- Change latency alarms from Average to p99 and use M of N evaluation.
- Set TreatMissingData deliberately on every alarm, with breaching for heartbeats.
- Create one composite alarm per user-facing service for paging, and route the rest to tickets.
- Save the Logs Insights queries your on-call engineers need, with narrow time ranges.
- If you need history beyond 15 months or finer than hourly after 63 days, set up a metric stream.