Route 53 is AWS's DNS service. It answers authoritative DNS queries for zones you host, registers domains, and runs health checks that change which answers it gives. That last capability is what makes it an availability tool rather than just a phone book: with the right records, Route 53 steers users away from a failed region without anyone touching a console. With the wrong ones it keeps sending traffic to a dead endpoint, or fails over so slowly that nobody notices it worked.

This article builds the model from the DNS query path up, explains alias records and every routing policy, shows exactly how health checks decide health and how Route 53 picks a record, walks through a multi-region failover setup with the CLI, and ends with timing math, failure modes and a checklist. Generic DNS failover and traffic-steering patterns across providers are compared in DNS failover and traffic steering; this page is about Route 53's own behaviour.

Advertisement

The query path and where Route 53 sits

When a browser looks up www.example.com, the operating system asks a recursive resolver, usually run by the ISP, a public DNS service or, inside a VPC, the Amazon-provided resolver. The resolver walks the hierarchy: the root servers point it at the .com servers, which return the NS records for example.com. If those NS records name Route 53's servers, the resolver asks Route 53, which is authoritative for the zone and returns the answer with a TTL. The resolver caches the answer for that TTL and serves it to every client that asks meanwhile.

Three facts follow that shape every Route 53 design. First, Route 53 only influences the answer at the moment a resolver asks; after that, caches hold it for the TTL. Second, the delegation at the registrar (the NS records in the parent zone) must match the hosted zone's name servers exactly, or queries never reach the zone you are editing. Third, health checks run continuously and independently of queries; a query reads the latest status rather than triggering a probe.

Resolving www.example.com with Route 53 health-checked failoverClientstub resolverRecursive resolvercaches for TTLRoot and .com TLDNS delegationRoute 53authoritative NSqueryPrimary: ALB us-east-1Secondary: ALB eu-west-1healthyelseHealth checkersmany locationsstatusprobe primaryHealth status is computed continuously; a query only reads it. Clients then keep the answer for up to the TTL
The recursive resolver finds Route 53 via NS delegation. Route 53 answers from precomputed health status, returning the primary record while it is healthy and the secondary otherwise; resolvers then cache the answer for the TTL.

Hosted zones and records

A public hosted zone answers queries from the internet. Each gets a set of four name servers that you configure at the registrar. A private hosted zone is associated with one or more VPCs and answers only queries from resources in them, through the VPC resolver; it is how db.internal.example.com can resolve to a private address in production and differently in staging. Networking fundamentals for those VPCs are in AWS VPC in depth.

Records are the usual DNS types, A, AAAA, CNAME, MX, TXT, NS, SOA, CAA and others, plus Route 53's own extension, the alias record. An alias looks like an A or AAAA record to clients but points at an AWS resource or another record in the same zone, and Route 53 resolves the target itself.

Alias recordCNAME record
Zone apex (example.com)AllowedNot allowed
TargetsSelected AWS resources (CloudFront, ELB, API Gateway, S3 website, Global Accelerator and more) or a record in the same zoneAny DNS name, anywhere
Query chargeNone for alias queries to AWS resourcesCharged
TTLNot settable for AWS resources; the resource's TTL is usedYou set it
Seen by dig asA or AAAACNAME
Tracks target IP changesYes, automaticallyVia the target's own DNS

Rule of thumb: whenever the target is a supported AWS resource, use an alias. It works at the apex, it costs nothing to query, and it follows the resource when its IP addresses change. A typical example is pointing both example.com and www at a distribution, as described in Amazon CloudFront in depth.

Advertisement

The eight routing policies

PolicyChooses the answer byTypical use
SimpleOne record, no selection logicA single resource
WeightedProportions you assign per recordCanary releases, gradual migrations
LatencyThe AWS Region with the lowest measured latency to the resolverMulti-region active-active
FailoverPrimary while healthy, else secondaryActive-passive disaster recovery
GeolocationContinent, country or US state of the userLegal or content localisation
GeoproximityDistance to resource locations, with an adjustable biasShifting load between locations
IP-basedThe client or resolver IP matched to CIDR collections you defineRouting known ISP or office ranges
Multivalue answerUp to eight healthy records, chosen at randomSimple client-side spreading

All but IP-based routing can be used in private hosted zones. Location for latency, geolocation and IP-based routing is the resolver's address unless the resolver forwards the client's subnet through the EDNS0 client-subnet extension, which is why users of a distant public resolver can be routed by the resolver's location rather than their own.

Policies combine by aliasing one record to another in the same zone, producing a tree: for example, latency records per region at the top, each aliasing a failover pair or a weighted set underneath. The Evaluate Target Health flag on each alias makes health propagate up the tree.

How health checks decide health

A Route 53 health check is not one probe. Health checkers in many locations probe the endpoint independently, and Route 53 aggregates their results. The documented rules:

  • Each checker applies the interval you choose, 30 seconds (standard) or 10 seconds (fast, extra charge), and a failure threshold: the number of consecutive failures or successes needed to flip that checker's view.
  • For HTTP and HTTPS checks, the TCP connection must succeed within four seconds and a 2xx or 3xx status must arrive within two seconds after connecting. TCP checks allow ten seconds to connect.
  • The endpoint is healthy if more than 18% of checkers report it healthy. A low threshold is deliberate: it stops a network problem between some checker locations and your endpoint from failing it globally.
  • String matching searches the first 5,120 bytes of the body, case-sensitively; compressed bodies are only decoded for gzip and deflate.
  • HTTPS checks do not validate certificates, so an expired certificate does not fail the check even though browsers reject it.
  • Checkers cannot reach private, local or non-routable addresses. For private endpoints, alarm on a CloudWatch metric and create a health check that monitors the alarm's data stream.

Two more types exist: calculated health checks combine child checks with AND, OR or at-least-N logic, and CloudWatch alarm checks, just mentioned, turn any metric into a health signal. A disabled health check is always treated as healthy; to force traffic away, invert it instead.

# Health check against the primary region's ALB: HTTPS, /healthz, 10 s interval, 3 failures
aws route53 create-health-check \
  --caller-reference "www-primary-2026-09-30" \
  --health-check-config '{
      "Type": "HTTPS",
      "FullyQualifiedDomainName": "<primary-alb-dns-name>",
      "Port": 443,
      "ResourcePath": "/healthz",
      "RequestInterval": 10,
      "FailureThreshold": 3,
      "EnableSNI": true
  }'

Design the endpoint behind /healthz to test what users need, a database round trip or a dependency flag, but make it cheap, because many checkers call it every few seconds.

How Route 53 picks a record

The record-selection rules are documented and short, and knowing them prevents most surprises. This pseudocode restates them for failover and weighted groups:

# Documented record-selection rules for one name + type + routing policy, as pseudocode
def healthy(rec):
    ok = True                                       # no health signal at all: always healthy
    if rec.health_check_id is not None:
        ok = health_status[rec.health_check_id]     # precomputed, not probed per query
    if rec.alias and rec.evaluate_target_health:
        ok = ok and target_is_healthy(rec.alias_target)  # both must pass when both are set
    return ok

def answer_failover(primary, secondary):
    if healthy(primary):
        return primary
    if healthy(secondary):
        return secondary
    return primary                                  # both unhealthy: primary is returned

def answer_weighted(records):
    live = [r for r in records if r.weight > 0 and healthy(r)]
    if not live:                                    # all nonzero weights down: try weight 0
        live = [r for r in records if r.weight == 0 and healthy(r)]
    if not live:
        live = records                              # none healthy: treat all as healthy
    return weighted_random_choice(live)

Note the two fail-open behaviours. If every record in a group is unhealthy, Route 53 treats them all as healthy and answers by policy, because returning nothing is worse than returning something. In a failover pair, if both are unhealthy, the primary is returned. And a record without any health check counts as permanently healthy, so a weighted group where one record lacks a check keeps receiving traffic for that record whatever happens.

Worked example: two-region active-passive failover

Goal: serve www.example.com from an Application Load Balancer in us-east-1 and fail over to one in eu-west-1. Create the health check above for the primary, then upsert a failover pair of alias records. The secondary has no health check of its own; its alias evaluates the ALB's target health.

{
  "Comment": "Active-passive failover for www.example.com",
  "Changes": [
    { "Action": "UPSERT",
      "ResourceRecordSet": {
        "Name": "www.example.com", "Type": "A",
        "SetIdentifier": "primary-us-east-1", "Failover": "PRIMARY",
        "HealthCheckId": "<primary-health-check-id>",
        "AliasTarget": { "HostedZoneId": "<alb-zone-id-us-east-1>",
                         "DNSName": "<primary-alb-dns-name>",
                         "EvaluateTargetHealth": true } } },
    { "Action": "UPSERT",
      "ResourceRecordSet": {
        "Name": "www.example.com", "Type": "A",
        "SetIdentifier": "secondary-eu-west-1", "Failover": "SECONDARY",
        "AliasTarget": { "HostedZoneId": "<alb-zone-id-eu-west-1>",
                         "DNSName": "<secondary-alb-dns-name>",
                         "EvaluateTargetHealth": true } } }
  ]
}
aws route53 change-resource-record-sets --hosted-zone-id <zone-id> --change-batch file://failover.json
aws route53 get-change --id <change-id>           # PENDING until propagated, then INSYNC
dig +short www.example.com @<one-of-the-zone-name-servers>

Why both a health check and EvaluateTargetHealth on the primary? Target health only knows whether the load balancer has healthy registered targets. The HTTPS check tests the full path users take, including listeners, security groups and the application's own dependencies. When an alias has both, the documentation states that both must pass, so either signal fails the primary. AWS generally recommends attaching health checks to non-alias records and relying on Evaluate Target Health for aliases; the combination is used here deliberately because the ALB alias is the only record in front of each region. The health check targets the load balancer's own DNS name, not www, because checking the name that is itself being failed over gives unpredictable results.

Now do the timing math, because it defines what failover actually means for users. With a 10-second interval and a threshold of 3, individual checkers need about 30 seconds of consecutive failures, and the aggregate flips once no more than 18% of checkers still see the endpoint healthy. After the flip, resolvers keep serving the cached primary answer for up to its TTL; alias records to an ALB use the load balancer's TTL, which you cannot raise or lower. Clients add their own caching: browsers, operating systems and runtimes such as the JVM may cache DNS independently. A realistic detection-plus-convergence window is therefore a minute or more for most users, with a tail of clients that hold on longer. Test it rather than trusting the arithmetic: block the health path on the primary and measure when traffic moves.

Operating Route 53

  • Pre-provision the failover. DNS query answering is designed for very high availability, but creating or changing records goes through the Route 53 API, a separate control plane. Build the records and health checks ahead of time so failover depends only on health status, never on an API call during an incident.
  • Rehearse. Invert the primary's health check to shift traffic deliberately, and measure how long clients take to move.
  • Alarm on health checks. Health check status is published to CloudWatch; the documentation places the SNS topics for these notifications in US East (N. Virginia).
  • Enable query logging on public zones to see which names and types resolvers ask for, useful when debugging delegation and geolocation.
  • Manage records as code with CloudFormation, Terraform or CDK, and review NS and apex records carefully: they are the ones that take a site down.
  • Plan TTL changes in advance. Lower a non-alias record's TTL at least one old TTL before a migration, so that caches have expired by the time you switch.

Failure modes

SymptomCauseFix
Edits have no effectRegistrar NS records point at a different or deleted hosted zoneCompare delegation with the zone's NS set
Failover never happensRecord has no health check, so it is always healthyAttach checks or set EvaluateTargetHealth
Traffic still reaches a dead regionAll records unhealthy, so all treated as healthyMake sure the fallback really is healthy
Health check flaps or fails oddlyIt targets the record's own name, which is itself being failed overCheck each endpoint by its own name or IP
String match always failsBody compressed with brotli, or string beyond 5,120 bytesServe gzip or plain text on the health path
Expired certificate, check still greenHTTPS checks do not validate certificatesMonitor certificate expiry separately
Failover took many minutesClient and runtime DNS caching beyond the TTLCap runtime DNS caches; test with real clients

Trade-offs

DNS-based failover is cheap and universal, but it is bounded by caching you do not control. When you need failover in seconds regardless of resolver behaviour, anycast front doors such as AWS Global Accelerator or a CDN route on stable IP addresses and move traffic without waiting for caches, at extra cost. Latency and geolocation policies are coarse because they see resolvers rather than users. Weighted routing is fine for gradual cutovers but does not guarantee exact splits, because resolvers cache and share answers across many clients.

What to do next

  1. Verify each production domain's delegation matches its hosted zone's four name servers.
  2. List records without health checks in every weighted, latency and failover group and decide whether each is intentional.
  3. Convert CNAMEs that target AWS resources to alias records, including the zone apex.
  4. Add a health check that exercises the real user path, with a CloudWatch alarm on its status.
  5. Run a failover drill by inverting the primary health check, and record the time until traffic moves.
  6. Move DNS records into infrastructure as code and require review for NS, apex and failover changes.
Key takeaway: Route 53 is authoritative DNS plus health-driven answer selection. Delegation must match the hosted zone, alias records should front AWS resources, and routing policies combine into trees through aliases with Evaluate Target Health. Health is decided by many checkers with an 18% rule and strict timeouts; records without checks are always healthy, and an all-unhealthy group fails open. Pre-build failover, rehearse it, and remember that TTLs and client caches, not Route 53, set how quickly users actually move.