Your dashboards say every service is healthy, yet customers cannot log in. Server-side metrics only see requests that arrive; they miss a broken DNS record, an expired certificate, a CDN serving a stale bundle, or a JavaScript error that stops the login button working. Synthetic monitoring closes that gap by acting like a user on a schedule, from outside your system, and failing loudly when the journey breaks.

Amazon CloudWatch Synthetics is AWS's managed version of that idea. You write a script, called a canary, that loads pages or calls APIs. AWS runs it on a schedule in a Lambda function, records screenshots, network archives and logs to S3, and publishes success and latency metrics you can alarm on. This article explains how a canary actually runs, how to write a reliable multi-step canary with Playwright, how to configure schedules, retries and alarms, how to reach private endpoints, what it costs, and which failure modes cause most canary noise.

Advertisement

What a canary is, mechanically

One canary run: schedule, Lambda with the Synthetics runtime layer, target, then artifacts, metrics and alarmsSchedulerate(5 minutes) / cronLambda functionyour script + runtime layerPublic endpointvia Lambda's internet egressPrivate endpoint in a VPCENIs, security groups, NATinvokeS3 artifactsscreenshots, HAR, logs, reportCloudWatch Logsper run, canaryRunIdMetricsCloudWatchSynthetics nsAlarmSuccessPercent < XSNS / EventBridge / incident toolEach run is a fresh Lambda invocation: no statesurvives between runs except what you write to S3.
The canary pipeline: a schedule invokes a Lambda function that runs your script with the Synthetics runtime, which uploads artifacts to S3, writes logs, and publishes metrics that drive alarms.

A canary is a Lambda function that CloudWatch Synthetics creates and manages for you. It contains your script plus a Synthetics runtime, which AWS documents as the Synthetics library code that calls your handler together with Lambda layers of bundled dependencies: a headless browser, the automation framework and helper libraries. Each scheduled run is an ordinary Lambda invocation. Nothing persists between runs except what the runtime uploads to S3, which matters when you design login flows and test data.

A run passes if your handler returns and fails if it throws or times out. After each run the runtime publishes metrics to the CloudWatchSynthetics namespace, writes the run's logs, and uploads artifacts such as screenshots, a HAR file of network requests and a run report to the S3 location you configured.

Runtimes and blueprints

A runtime version name follows the pattern syn-{language}-{framework}-{major}.{minor}. CloudWatch Synthetics currently offers Node.js with Puppeteer, Node.js with Playwright, Python with Selenium WebDriver, and Java runtimes. At the time of writing the newest Playwright runtime is syn-nodejs-playwright-8.0, built on the Lambda Node.js 22.x runtime with Playwright 1.61.1, Chromium and Firefox. From syn-nodejs-playwright-5.1 the library moved from @amzn/synthetics-playwright to @aws/synthetics-playwright, and the legacy name is slated for removal. Playwright runtimes from 3.0 can run a canary in Firefox as well as Chrome, and 7.0 added multi-location canaries that run the same canary across several Regions from one definition.

AWS deprecates runtime versions on a published schedule. Pin the runtime explicitly, read the support policy page when you create a canary, and plan runtime upgrades like any other dependency upgrade, with a dry run first.

For common checks you do not have to write code from scratch. The console offers blueprints for a heartbeat monitor (load a URL, save a screenshot), an API canary, a broken-link checker, visual monitoring that compares screenshots against a baseline, a canary recorder that captures clicks from a browser extension, and a GUI workflow builder.

Advertisement

Writing a multi-step Playwright canary

The most useful canaries follow a real user journey and split it into named steps. The Playwright runtime's library exposes launch, newPage, executeStep, getDefaultLaunchOptions and close. newPage sets up the network capture that produces the HAR file. executeStep takes screenshots around the step, records its result in the run report, and publishes SuccessPercent and Duration with a StepName dimension, so you can see which step broke and how long each one takes. Here is a canary for a shop: load the home page, log in as a test user, search, and check a JSON API.

// index.js -- Node.js Playwright runtime (syn-nodejs-playwright-5.1 or later uses this namespace)
const { synthetics } = require('@aws/synthetics-playwright');
const { SecretsManagerClient, GetSecretValueCommand } = require('@aws-sdk/client-secrets-manager');

const BASE = process.env.BASE_URL;                       // e.g. https://shop.example.com

async function testUser() {
  const sm = new SecretsManagerClient({});
  const out = await sm.send(new GetSecretValueCommand({ SecretId: 'canary/shop-test-user' }));
  return JSON.parse(out.SecretString);                   // { "email": "...", "password": "..." }
}

exports.handler = async () => {
  const browser = await synthetics.launch();
  try {
    const page = await synthetics.newPage(browser);
    const user = await testUser();

    await synthetics.executeStep('home', async () => {
      const resp = await page.goto(BASE, { waitUntil: 'load', timeout: 30000 });
      if (!resp || resp.status() >= 400) throw new Error(`home returned ${resp && resp.status()}`);
    });

    await synthetics.executeStep('login', async () => {
      await page.getByLabel('Email').fill(user.email);
      await page.getByLabel('Password').fill(user.password);
      await page.getByRole('button', { name: 'Sign in' }).click();
      await page.getByTestId('account-menu').waitFor({ timeout: 10000 });
    });

    await synthetics.executeStep('search', async () => {
      await page.getByRole('searchbox').fill('canary-test-sku');
      await page.keyboard.press('Enter');
      await page.getByTestId('result-card').first().waitFor({ timeout: 10000 });
    });

    await synthetics.executeStep('price-api', async () => {
      const r = await fetch(`${BASE}/api/v1/price/canary-test-sku`, { signal: AbortSignal.timeout(5000) });
      if (r.status !== 200) throw new Error(`price API ${r.status}`);
      const body = await r.json();
      if (typeof body.cents !== 'number') throw new Error('price API returned no cents field');
    }, { screenshotOnStepStart: false, screenshotOnStepSuccess: false });
  } finally {
    await synthetics.close();
  }
};

Several choices in this script prevent noise later. Selectors use roles, labels and test IDs, not CSS paths that change with every redesign. Each wait has an explicit timeout shorter than the canary's own, so a hang fails the right step with a useful message rather than timing out the whole run. Credentials come from Secrets Manager at run time rather than from environment variables, which are visible to anyone who can read the canary's configuration. The API step disables screenshots, since there is nothing to see. The finally block closes the browser, which the library documentation recommends.

Write artifacts only to /tmp, the only writable directory in the Lambda environment; the runtime uploads what it finds there. Package dependencies with the folder layout your runtime version documents; layouts changed between major versions, and a wrong layout produces a handler-not-found error on the first run.

Schedules, timeouts and retries

A canary runs on a rate expression from rate(1 minute) to rate(1 hour), or on a cron expression. The special value rate(0 minute) runs it once when started, which is useful for deployment smoke tests triggered from a pipeline. If you do not set a timeout, the canary's frequency is used as its timeout, up to a maximum of 900 seconds. Set it explicitly and well below the interval so runs never overlap.

Retries are configured with MaxRetries between 0 and 2. A retry re-runs a failed canary immediately, and the SuccessPercentWithRetries metric reports success after all attempts, while SuccessPercent and RetryCount keep the first-attempt picture visible. The API reference notes that with MaxRetries = 2 the run timeout must be under 600 seconds. Retries hide transient network blips, but they also hide intermittent real failures, so alarm on one metric and watch the other on a dashboard.

Dry runs, supported from syn-nodejs-playwright-2.0 in the Playwright line, execute a changed script once without replacing the live version, and publish their own SuccessPercentDryRun and DurationDryRun metrics.

Infrastructure as code and alarm design

Treat canaries as production code: store scripts in the repository, package them in CI, and deploy them with CloudFormation, CDK or Terraform. A minimal CloudFormation definition with an alarm looks like this:

ShopCanary:
  Type: AWS::Synthetics::Canary
  Properties:
    Name: shop-login-search
    RuntimeVersion: syn-nodejs-playwright-8.0       # pin, then upgrade deliberately
    ExecutionRoleArn: !GetAtt CanaryRole.Arn
    ArtifactS3Location: !Sub s3://${CanaryArtifacts}/shop-login-search
    Code:
      Handler: index.handler
      S3Bucket: !Ref CanaryCodeBucket
      S3Key: canaries/shop-login-search.zip
    Schedule:
      Expression: rate(5 minutes)
      RetryConfig:
        MaxRetries: 1
    RunConfig:
      TimeoutInSeconds: 120
      EnvironmentVariables:
        BASE_URL: https://shop.example.com
    SuccessRetentionPeriod: 7
    FailureRetentionPeriod: 30
    StartCanaryAfterCreation: true

ShopCanaryAlarm:
  Type: AWS::CloudWatch::Alarm
  Properties:
    Namespace: CloudWatchSynthetics
    MetricName: SuccessPercent
    Dimensions:
      - { Name: CanaryName, Value: shop-login-search }
    Statistic: Average
    Period: 300
    EvaluationPeriods: 3
    DatapointsToAlarm: 2
    Threshold: 100
    ComparisonOperator: LessThanThreshold
    TreatMissingData: breaching
    AlarmActions: [ !Ref OnCallTopic ]

The alarm is where most teams go wrong. A canary at rate(5 minutes) produces one data point per five-minute period, so a 300-second period with DatapointsToAlarm: 2 out of three means two failed runs in fifteen minutes page someone, and a single flaky run does not. TreatMissingData: breaching matters: if the canary stops running, because of a deleted role, an exhausted Lambda concurrency limit or a deprecated runtime, missing data should alarm rather than look healthy. The alarm deliberately watches SuccessPercent, the first-attempt metric, so the retry adds evidence but does not hide a failing first run; alarm on SuccessPercentWithRetries instead if you want retries to suppress pages. At rate(1 minute) use a one-minute period and an M-of-N rule such as three of five.

Send the alarm to an SNS topic or use EventBridge rules on the alarm state change to fan out to chat and incident tooling; SNS covers subscription filtering and delivery retries. Add a second, lower-urgency alarm on step Duration to catch slow degradation before it becomes failure.

Reaching private endpoints: VPC canaries

Canaries run outside your VPC by default and can reach only public endpoints. To test an internal API, a private load balancer or a database-backed health page, attach the canary to VPC subnets and security groups, as with any VPC-attached Lambda function. The function then gets elastic network interfaces in your subnets. Two consequences follow. A VPC canary has no internet access unless its subnets route through a NAT gateway, and it also needs a route to S3 and CloudWatch for its artifacts and metrics, through the NAT gateway or VPC endpoints. And because outbound traffic leaves through the NAT gateway's elastic IP, a VPC canary gives you a stable source address that a partner's firewall or a WAF rule can allow-list, which a non-VPC canary cannot.

The execution role needs permission to write artifacts to the S3 bucket, create log streams and put log events, and publish metrics; restrict cloudwatch:PutMetricData with a condition on the CloudWatchSynthetics namespace. VPC canaries also need the usual network-interface permissions for Lambda. Add read access to exactly the secrets the script uses, and nothing else.

Worked example: what a canary fleet costs

Synthetics bills per canary run, with a monthly free tier of runs, plus the Lambda, S3, CloudWatch Logs and alarm usage it creates. At the time of writing the published price starts at $0.0012 per run in the cheapest Regions; check the CloudWatch pricing page for your Region before budgeting.

A canary at rate(1 minute) runs 60 x 24 x 30 = 43,200 times in a 30-day month, which is about $51.84 in run charges alone. The same canary at rate(5 minutes) runs 8,640 times, about $10.37. Ten journeys at one minute in three Regions is 1,296,000 runs, roughly $1,555 a month before Lambda, storage and logs. Retries add runs on failure, and artifact storage grows with every screenshot, so set SuccessRetentionPeriod short and FailureRetentionPeriod long, and add an S3 lifecycle rule as a backstop.

The arithmetic suggests a tiered design. Run cheap API or heartbeat checks every minute on the endpoints that define availability, run full browser journeys every five or ten minutes, and run broken-link and visual checks hourly.

Failure modes

  • Brittle selectors. CSS paths and text matches break on harmless UI changes and train people to ignore the alarm. Use roles, labels and dedicated test IDs.
  • Canaries that change production data. A checkout canary that places real orders pollutes revenue and stock. Use a test tenant, a test SKU, or stop before payment, and tag canary traffic so analytics can exclude it.
  • Shared test accounts. Lockouts after repeated failed logins, MFA prompts and expired passwords look like outages. Give canaries dedicated accounts exempt from MFA by network or by policy, and rotate their secrets automatically.
  • Silent canaries. A stopped canary, a missing permission or a deprecated runtime produces no data, not failures. Alarm on missing data and review runtime deprecation notices.
  • Secrets in environment variables or screenshots. Load credentials at run time and avoid screenshots of pages that display tokens or personal data.
  • One Region only. A canary in the same Region as the application shares its failure domain. Run critical journeys from at least one other Region.

Trade-offs

Synthetics gives you managed scheduling, artifacts and metrics inside your AWS account and IAM model, with no agents to run. In exchange, locations are limited to AWS Regions rather than residential or mobile networks, browser coverage depends on the runtime, and per-run pricing makes high-frequency browser checks across many Regions expensive. Real user monitoring complements rather than replaces canaries: it measures what users actually experience, but only when users are present and only for paths they take, while a canary probes a known path at 3 a.m. when traffic is zero.

For simple reachability, a Route 53 or load balancer health check is cheaper. Use canaries for multi-step journeys, content assertions and API contracts that health checks cannot express. For the rest of the monitoring stack, alarms, metric math and dashboards, see CloudWatch in depth, and put the canary's API checks next to the API Gateway stages they exercise.

What to do next

  1. List the three to five user journeys that define availability for your product, such as sign-in, search and checkout up to payment.
  2. Create a test tenant, a test SKU and dedicated test accounts with secrets in Secrets Manager.
  3. Write one Playwright canary per journey with named executeStep steps, role-based selectors and explicit step timeouts; pin the runtime version.
  4. Deploy canaries with infrastructure as code, with a dry run in the pipeline before every change.
  5. Alarm on SuccessPercent with an M-of-N rule and missing data treated as breaching; add a lower-urgency alarm on step duration.
  6. Add a VPC canary for one internal dependency and confirm its NAT or VPC endpoint routes for S3 and CloudWatch.
  7. Set artifact retention and an S3 lifecycle rule, then compute the monthly cost per canary with the run arithmetic above.
  8. Subscribe to runtime deprecation notices and schedule upgrades quarterly.
Key takeaway: CloudWatch Synthetics runs your scripts as managed Lambda functions on a schedule, records screenshots, HAR files and logs to S3, and publishes SuccessPercent and Duration metrics, per step when you use executeStep. Write canaries for real user journeys with stable selectors, explicit timeouts and secrets loaded at run time, pin and upgrade runtimes deliberately, and alarm with an M-of-N rule that treats missing data as failure. Use VPC canaries for private endpoints and a stable egress IP, and budget from the run count: a one-minute canary is about 43,200 runs a month. Tier frequency by cost so cheap checks catch outages fast and full browser journeys catch functional breakage.