Route 53 is AWS's DNS service. It answers authoritative DNS queries for zones you host, registers domains, and runs health checks that change which answers it gives. That last capability is what makes it an availability tool rather than just a phone book: with the right records, Route 53 steers users away from a failed region without anyone touching a console. With the wrong ones it keeps sending traffic to a dead endpoint, or fails over so slowly that nobody notices it worked.
This article builds the model from the DNS query path up, explains alias records and every routing policy, shows exactly how health checks decide health and how Route 53 picks a record, walks through a multi-region failover setup with the CLI, and ends with timing math, failure modes and a checklist. Generic DNS failover and traffic-steering patterns across providers are compared in DNS failover and traffic steering; this page is about Route 53's own behaviour.
The query path and where Route 53 sits
When a browser looks up www.example.com, the operating system asks a recursive resolver, usually run by the ISP, a public DNS service or, inside a VPC, the Amazon-provided resolver. The resolver walks the hierarchy: the root servers point it at the .com servers, which return the NS records for example.com. If those NS records name Route 53's servers, the resolver asks Route 53, which is authoritative for the zone and returns the answer with a TTL. The resolver caches the answer for that TTL and serves it to every client that asks meanwhile.
Three facts follow that shape every Route 53 design. First, Route 53 only influences the answer at the moment a resolver asks; after that, caches hold it for the TTL. Second, the delegation at the registrar (the NS records in the parent zone) must match the hosted zone's name servers exactly, or queries never reach the zone you are editing. Third, health checks run continuously and independently of queries; a query reads the latest status rather than triggering a probe.
Hosted zones and records
A public hosted zone answers queries from the internet. Each gets a set of four name servers that you configure at the registrar. A private hosted zone is associated with one or more VPCs and answers only queries from resources in them, through the VPC resolver; it is how db.internal.example.com can resolve to a private address in production and differently in staging. Networking fundamentals for those VPCs are in AWS VPC in depth.
Records are the usual DNS types, A, AAAA, CNAME, MX, TXT, NS, SOA, CAA and others, plus Route 53's own extension, the alias record. An alias looks like an A or AAAA record to clients but points at an AWS resource or another record in the same zone, and Route 53 resolves the target itself.
| Alias record | CNAME record | |
|---|---|---|
| Zone apex (example.com) | Allowed | Not allowed |
| Targets | Selected AWS resources (CloudFront, ELB, API Gateway, S3 website, Global Accelerator and more) or a record in the same zone | Any DNS name, anywhere |
| Query charge | None for alias queries to AWS resources | Charged |
| TTL | Not settable for AWS resources; the resource's TTL is used | You set it |
| Seen by dig as | A or AAAA | CNAME |
| Tracks target IP changes | Yes, automatically | Via the target's own DNS |
Rule of thumb: whenever the target is a supported AWS resource, use an alias. It works at the apex, it costs nothing to query, and it follows the resource when its IP addresses change. A typical example is pointing both example.com and www at a distribution, as described in Amazon CloudFront in depth.
The eight routing policies
| Policy | Chooses the answer by | Typical use |
|---|---|---|
| Simple | One record, no selection logic | A single resource |
| Weighted | Proportions you assign per record | Canary releases, gradual migrations |
| Latency | The AWS Region with the lowest measured latency to the resolver | Multi-region active-active |
| Failover | Primary while healthy, else secondary | Active-passive disaster recovery |
| Geolocation | Continent, country or US state of the user | Legal or content localisation |
| Geoproximity | Distance to resource locations, with an adjustable bias | Shifting load between locations |
| IP-based | The client or resolver IP matched to CIDR collections you define | Routing known ISP or office ranges |
| Multivalue answer | Up to eight healthy records, chosen at random | Simple client-side spreading |
All but IP-based routing can be used in private hosted zones. Location for latency, geolocation and IP-based routing is the resolver's address unless the resolver forwards the client's subnet through the EDNS0 client-subnet extension, which is why users of a distant public resolver can be routed by the resolver's location rather than their own.
Policies combine by aliasing one record to another in the same zone, producing a tree: for example, latency records per region at the top, each aliasing a failover pair or a weighted set underneath. The Evaluate Target Health flag on each alias makes health propagate up the tree.
How health checks decide health
A Route 53 health check is not one probe. Health checkers in many locations probe the endpoint independently, and Route 53 aggregates their results. The documented rules:
- Each checker applies the interval you choose, 30 seconds (standard) or 10 seconds (fast, extra charge), and a failure threshold: the number of consecutive failures or successes needed to flip that checker's view.
- For HTTP and HTTPS checks, the TCP connection must succeed within four seconds and a 2xx or 3xx status must arrive within two seconds after connecting. TCP checks allow ten seconds to connect.
- The endpoint is healthy if more than 18% of checkers report it healthy. A low threshold is deliberate: it stops a network problem between some checker locations and your endpoint from failing it globally.
- String matching searches the first 5,120 bytes of the body, case-sensitively; compressed bodies are only decoded for gzip and deflate.
- HTTPS checks do not validate certificates, so an expired certificate does not fail the check even though browsers reject it.
- Checkers cannot reach private, local or non-routable addresses. For private endpoints, alarm on a CloudWatch metric and create a health check that monitors the alarm's data stream.
Two more types exist: calculated health checks combine child checks with AND, OR or at-least-N logic, and CloudWatch alarm checks, just mentioned, turn any metric into a health signal. A disabled health check is always treated as healthy; to force traffic away, invert it instead.
# Health check against the primary region's ALB: HTTPS, /healthz, 10 s interval, 3 failures
aws route53 create-health-check \
--caller-reference "www-primary-2026-09-30" \
--health-check-config '{
"Type": "HTTPS",
"FullyQualifiedDomainName": "<primary-alb-dns-name>",
"Port": 443,
"ResourcePath": "/healthz",
"RequestInterval": 10,
"FailureThreshold": 3,
"EnableSNI": true
}'Design the endpoint behind /healthz to test what users need, a database round trip or a dependency flag, but make it cheap, because many checkers call it every few seconds.
How Route 53 picks a record
The record-selection rules are documented and short, and knowing them prevents most surprises. This pseudocode restates them for failover and weighted groups:
# Documented record-selection rules for one name + type + routing policy, as pseudocode
def healthy(rec):
ok = True # no health signal at all: always healthy
if rec.health_check_id is not None:
ok = health_status[rec.health_check_id] # precomputed, not probed per query
if rec.alias and rec.evaluate_target_health:
ok = ok and target_is_healthy(rec.alias_target) # both must pass when both are set
return ok
def answer_failover(primary, secondary):
if healthy(primary):
return primary
if healthy(secondary):
return secondary
return primary # both unhealthy: primary is returned
def answer_weighted(records):
live = [r for r in records if r.weight > 0 and healthy(r)]
if not live: # all nonzero weights down: try weight 0
live = [r for r in records if r.weight == 0 and healthy(r)]
if not live:
live = records # none healthy: treat all as healthy
return weighted_random_choice(live)Note the two fail-open behaviours. If every record in a group is unhealthy, Route 53 treats them all as healthy and answers by policy, because returning nothing is worse than returning something. In a failover pair, if both are unhealthy, the primary is returned. And a record without any health check counts as permanently healthy, so a weighted group where one record lacks a check keeps receiving traffic for that record whatever happens.
Worked example: two-region active-passive failover
Goal: serve www.example.com from an Application Load Balancer in us-east-1 and fail over to one in eu-west-1. Create the health check above for the primary, then upsert a failover pair of alias records. The secondary has no health check of its own; its alias evaluates the ALB's target health.
{
"Comment": "Active-passive failover for www.example.com",
"Changes": [
{ "Action": "UPSERT",
"ResourceRecordSet": {
"Name": "www.example.com", "Type": "A",
"SetIdentifier": "primary-us-east-1", "Failover": "PRIMARY",
"HealthCheckId": "<primary-health-check-id>",
"AliasTarget": { "HostedZoneId": "<alb-zone-id-us-east-1>",
"DNSName": "<primary-alb-dns-name>",
"EvaluateTargetHealth": true } } },
{ "Action": "UPSERT",
"ResourceRecordSet": {
"Name": "www.example.com", "Type": "A",
"SetIdentifier": "secondary-eu-west-1", "Failover": "SECONDARY",
"AliasTarget": { "HostedZoneId": "<alb-zone-id-eu-west-1>",
"DNSName": "<secondary-alb-dns-name>",
"EvaluateTargetHealth": true } } }
]
}aws route53 change-resource-record-sets --hosted-zone-id <zone-id> --change-batch file://failover.json
aws route53 get-change --id <change-id> # PENDING until propagated, then INSYNC
dig +short www.example.com @<one-of-the-zone-name-servers>Why both a health check and EvaluateTargetHealth on the primary? Target health only knows whether the load balancer has healthy registered targets. The HTTPS check tests the full path users take, including listeners, security groups and the application's own dependencies. When an alias has both, the documentation states that both must pass, so either signal fails the primary. AWS generally recommends attaching health checks to non-alias records and relying on Evaluate Target Health for aliases; the combination is used here deliberately because the ALB alias is the only record in front of each region. The health check targets the load balancer's own DNS name, not www, because checking the name that is itself being failed over gives unpredictable results.
Now do the timing math, because it defines what failover actually means for users. With a 10-second interval and a threshold of 3, individual checkers need about 30 seconds of consecutive failures, and the aggregate flips once no more than 18% of checkers still see the endpoint healthy. After the flip, resolvers keep serving the cached primary answer for up to its TTL; alias records to an ALB use the load balancer's TTL, which you cannot raise or lower. Clients add their own caching: browsers, operating systems and runtimes such as the JVM may cache DNS independently. A realistic detection-plus-convergence window is therefore a minute or more for most users, with a tail of clients that hold on longer. Test it rather than trusting the arithmetic: block the health path on the primary and measure when traffic moves.
Operating Route 53
- Pre-provision the failover. DNS query answering is designed for very high availability, but creating or changing records goes through the Route 53 API, a separate control plane. Build the records and health checks ahead of time so failover depends only on health status, never on an API call during an incident.
- Rehearse. Invert the primary's health check to shift traffic deliberately, and measure how long clients take to move.
- Alarm on health checks. Health check status is published to CloudWatch; the documentation places the SNS topics for these notifications in US East (N. Virginia).
- Enable query logging on public zones to see which names and types resolvers ask for, useful when debugging delegation and geolocation.
- Manage records as code with CloudFormation, Terraform or CDK, and review NS and apex records carefully: they are the ones that take a site down.
- Plan TTL changes in advance. Lower a non-alias record's TTL at least one old TTL before a migration, so that caches have expired by the time you switch.
Failure modes
| Symptom | Cause | Fix |
|---|---|---|
| Edits have no effect | Registrar NS records point at a different or deleted hosted zone | Compare delegation with the zone's NS set |
| Failover never happens | Record has no health check, so it is always healthy | Attach checks or set EvaluateTargetHealth |
| Traffic still reaches a dead region | All records unhealthy, so all treated as healthy | Make sure the fallback really is healthy |
| Health check flaps or fails oddly | It targets the record's own name, which is itself being failed over | Check each endpoint by its own name or IP |
| String match always fails | Body compressed with brotli, or string beyond 5,120 bytes | Serve gzip or plain text on the health path |
| Expired certificate, check still green | HTTPS checks do not validate certificates | Monitor certificate expiry separately |
| Failover took many minutes | Client and runtime DNS caching beyond the TTL | Cap runtime DNS caches; test with real clients |
Trade-offs
DNS-based failover is cheap and universal, but it is bounded by caching you do not control. When you need failover in seconds regardless of resolver behaviour, anycast front doors such as AWS Global Accelerator or a CDN route on stable IP addresses and move traffic without waiting for caches, at extra cost. Latency and geolocation policies are coarse because they see resolvers rather than users. Weighted routing is fine for gradual cutovers but does not guarantee exact splits, because resolvers cache and share answers across many clients.
What to do next
- Verify each production domain's delegation matches its hosted zone's four name servers.
- List records without health checks in every weighted, latency and failover group and decide whether each is intentional.
- Convert CNAMEs that target AWS resources to alias records, including the zone apex.
- Add a health check that exercises the real user path, with a CloudWatch alarm on its status.
- Run a failover drill by inverting the primary health check, and record the time until traffic moves.
- Move DNS records into infrastructure as code and require review for NS, apex and failover changes.