Most multi-region designs on AWS fail over with DNS: change the answer, wait for resolvers and clients to notice. That wait is the weakness. Resolvers and client libraries cache answers, sometimes well beyond the record's TTL, and some clients never re-resolve at all. Some clients cannot use DNS names in the first place: firewall allow-lists, IoT devices with hard-coded addresses, and game or VoIP clients that want a fixed IP.

AWS Global Accelerator takes a different approach. It gives your application static anycast IP addresses announced from AWS edge locations worldwide. Clients connect to the nearest edge, and AWS carries traffic over its own backbone to healthy endpoints in one or more Regions. Failover happens behind the IP address, not in DNS. This article explains how that works, what the traffic controls actually do, where the sharp edges are, and how it compares to Route 53 and CloudFront.

Advertisement

The model: anycast in, backbone across

An accelerator gets two static IPv4 addresses (four with dual-stack: two IPv4 and two IPv6), or addresses from your own range if you bring your own IPv4 addresses. The same addresses are announced by BGP from many edge locations, so the internet routes each client to a nearby edge. That is anycast: one address, many places.

For TCP, Global Accelerator terminates the client's connection at the edge and opens a separate connection to your endpoint almost at once. The handshake round trip is to a nearby edge, not a distant Region, and the long leg runs over the AWS network instead of the public internet. For UDP, packets are forwarded. The addresses stay with the accelerator for its lifetime, even while it is disabled, and are lost only when you delete it, so restrict deletion with IAM.

Global Accelerator: anycast IPs at the edge, then the AWS backbone to regional endpointsClient in ParisClient in TokyoEdge (Europe)anycast IPs, TCP endsEdge (Asia)anycast IPs, TCP endsBGPBGPListenerTCP 443, affinitybackbonebackboneGroup eu-west-1dial 100%, ALB w128Group ap-northeast-1dial 100%, ALB w128Health checksELB health or GA checksClients always dial the same two IPs. Health, traffic dials and weights decide which region and endpoint serve each new connection.
Clients reach the nearest edge via anycast. The listener hands each new connection to an endpoint group chosen by proximity, health and traffic dial, then to an endpoint by weight.

Standard and custom routing accelerators

A standard accelerator sends traffic to the best healthy endpoint by client location, endpoint health, traffic dials and weights. Endpoints are Application Load Balancers, Network Load Balancers, EC2 instances or Elastic IP addresses. This is the one for web APIs, multi-region failover and blue-green shifts.

A custom routing accelerator maps listener ports deterministically onto specific EC2 instance IP and port destinations in VPC subnets. Your own matchmaking logic tells a client "connect to accelerator IP, port 10017", and that always reaches one particular game or media server. There are no health checks and no failover, because you chose the destination, and it supports IPv4 only. The rest of this article is about standard accelerators.

Advertisement

Listeners, endpoint groups and endpoints

Three objects make up a standard accelerator. A listener accepts TCP or UDP on port ranges and sets client affinity. An endpoint group belongs to one listener and one Region; a listener has at most one group per Region. An endpoint is a load balancer, instance or Elastic IP inside that group, with a weight.

Port overrides let the listener accept 443 while endpoints receive a different port, such as 1443. Health-check ports cannot be overridden. The control-plane API is called in us-west-2 regardless of where endpoints run, which catches people whose automation assumes the local Region.

# The Global Accelerator control-plane API is called in us-west-2, whatever Regions the endpoints are in
aws globalaccelerator create-accelerator --name web-prod --ip-address-type IPV4 \
    --region us-west-2

aws globalaccelerator create-listener --accelerator-arn "$ACC_ARN" \
    --protocol TCP --port-ranges FromPort=443,ToPort=443 \
    --client-affinity NONE --region us-west-2

aws globalaccelerator create-endpoint-group --listener-arn "$LSN_ARN" \
    --endpoint-group-region eu-west-1 \
    --traffic-dial-percentage 100 \
    --endpoint-configurations EndpointId="$ALB_EU_ARN",Weight=128,ClientIPPreservationEnabled=true \
    --region us-west-2

The same in Terraform:

resource "aws_globalaccelerator_accelerator" "web" {
  name            = "web-prod"
  ip_address_type = "IPV4"
  enabled         = true
}

resource "aws_globalaccelerator_listener" "https" {
  accelerator_arn = aws_globalaccelerator_accelerator.web.id
  protocol        = "TCP"
  client_affinity = "NONE"
  port_range {
    from_port = 443
    to_port   = 443
  }
}

resource "aws_globalaccelerator_endpoint_group" "tokyo" {
  listener_arn            = aws_globalaccelerator_listener.https.id
  endpoint_group_region   = "ap-northeast-1"
  traffic_dial_percentage = 100
  endpoint_configuration {
    endpoint_id                    = aws_lb.tokyo.arn
    weight                         = 128
    client_ip_preservation_enabled = true
  }
}

Health checks and failover

Global Accelerator sends new connections only to healthy endpoints. How it decides depends on endpoint type. For load balancers, it uses Elastic Load Balancing health: an ALB counts as healthy when every target group has at least one healthy target, and an NLB when at least one Availability Zone is healthy. Health-check settings on the endpoint group do not apply to load balancers; configure them on the target groups.

For EC2 and Elastic IP endpoints, the endpoint group's own settings apply: protocol TCP, HTTP or HTTPS (default TCP), path (default /), port (default the listener's first port), an interval of 10 or 30 seconds (default 30) and a threshold of 1 to 10 consecutive results (default 3). UDP listeners are health-checked over TCP on the listener port, so the endpoint needs a TCP server there or it is marked unhealthy. Security groups must allow the Route 53 health-checker ranges.

Detection time is roughly interval times threshold: about 90 seconds with defaults, about 30 seconds at 10 seconds and 3. When a Region has no healthy endpoints, new connections go to the next-best Region. If no endpoint anywhere is healthy, Global Accelerator fails open and sends traffic to all endpoints, betting that a health check is wrong rather than everything being down.

Traffic dials and weights

Two controls shape traffic. The traffic dial on an endpoint group, from 0 to 100 percent (default 100), limits how much of the traffic already routed to that Region it accepts; the rest spills to other Regions. It is not a global share. With a dial of 50, a group that would have received 100 connections by proximity accepts 50 and the other 50 go elsewhere. A dial of 0 drains a Region for maintenance.

Weights split a group's traffic between its endpoints, from 0 to 255 (default 128); each endpoint receives its weight divided by the sum. Weights 192 and 64 give a 75/25 split, useful for canarying a new stack in the same Region. Global Accelerator may override weights in limited cases to protect availability, for example to avoid connection collisions with client IP preservation.

import boto3

# The client must target us-west-2 for Global Accelerator API calls.
ga = boto3.client("globalaccelerator", region_name="us-west-2")

def drain_region(endpoint_group_arn: str, steps=(75, 50, 25, 0)):
    """Dial a Region down in stages; existing flows finish, new ones go elsewhere."""
    for pct in steps:
        ga.update_endpoint_group(
            EndpointGroupArn=endpoint_group_arn,
            TrafficDialPercentage=float(pct),
        )
        wait_and_check_error_rates()        # your own SLO check between steps

Changes affect new connections only. Established flows stay on their endpoint until they close or hit the idle timeout, even if the endpoint is marked unhealthy or removed. Draining is therefore gradual, which is good for users and bad if you expected an instant cut.

Client affinity, client IPs and timeouts

Client affinity NONE hashes the five-tuple (source IP and port, destination IP and port, protocol), so two connections from one client can land on different endpoints. SOURCE_IP hashes source and destination IP only, keeping a client on one endpoint while it stays healthy; use it for stateful protocols that span connections. It does not replace real session storage.

With client IP preservation, endpoints see the real client address rather than a Global Accelerator address, so security groups, logs and rate limits work per client. Custom routing always preserves it; standard accelerators support it for ALB and EC2 endpoints, and for NLBs that have security groups and do not terminate TLS. AWS recommends not also sending direct internet traffic to preserved endpoints, and disabling cross-zone load balancing on NLB endpoints, to avoid connection collisions that delay TCP setup.

Idle timeouts are fixed: 340 seconds for TCP and 30 seconds for UDP. TCP keep-alive probes do not count; at least one byte of data must flow. Long-lived idle connections, such as WebSockets or database sessions, need application-level heartbeats more often than every 340 seconds.

Worked example: two-region API with fast failover

An API runs behind ALBs in eu-west-1 and ap-northeast-1, and partners allow-list its IPs. The team creates one accelerator, a TCP 443 listener with affinity NONE, and one endpoint group per Region holding that Region's ALB at weight 128 with client IP preservation on. Partners allow-list the two static IPs once. DNS points the API name at the accelerator's DNS name, but clients that hard-code the IPs work too.

Proximity sends European clients to Ireland and Asian clients to Tokyo. When the Tokyo ALB's target groups lose all healthy targets, ELB health marks the endpoint unhealthy, and new Asian connections go to Ireland over the backbone, with no DNS change and no client action. Existing Tokyo connections break as the targets fail; clients reconnect, and the new connections land in Ireland. Latency for Asian users rises for the duration, so Ireland must have capacity for both regions' peak, or autoscaling must cover the gap.

For a planned Tokyo deployment, the team drains with dials from 100 to 0 in steps, watches error rates, deploys, and dials back up. To canary, they add a second ALB to the Tokyo group at weight 13 beside the old one at 115, about 10 percent.

Global Accelerator, Route 53 or CloudFront

OptionFailover mechanismStatic IPsBest for
Global AcceleratorBehind anycast IPs, new connections moveYes, two per acceleratorTCP/UDP apps, fast regional failover, allow-lists
Route 53 failover or latency recordsChanges DNS answers; resolvers and clients cacheNoCheap multi-region routing where DNS caching is acceptable
CloudFrontEdge caching, origin failover per requestNoHTTP content, caching, edge compute

They combine. CloudFront in front of a Global Accelerator endpoint is unusual; Route 53 alias records pointing at an accelerator are common. See Route 53 in depth, CloudFront in depth and multi-region architecture on AWS. For security groups and subnets behind the endpoints, see VPC design.

Failure modes

SymptomLikely causeFirst response
Traffic still reaches an unhealthy endpointEstablished flows stay until idle timeoutExpect gradual drain; break connections at the app if needed
Idle WebSockets drop after about six minutes340 s TCP idle timeout; keep-alives ignoredSend application heartbeats
Every endpoint gets traffic during an outageNo healthy endpoints: fail-openFix health checks; do not rely on it to shed load
UDP endpoint always unhealthyNo TCP server on the health-check portRun a TCP health responder
Large UDP or fragmented TCP failsTCP fragments dropped at the edgeAllow ICMP for path MTU discovery; avoid fragmentation
Ping looks great, app is slowICMP answered at the edgeMeasure with real TCP or HTTP requests
Slow TCP setup with preserved client IPsConnection collisionsDisable NLB cross-zone; stop direct internet traffic to endpoints

What to do next

  1. Decide which clients need static IPs or faster-than-DNS failover; if none do, Route 53 may be enough.
  2. Create a standard accelerator with one endpoint group per Region, calling the API in us-west-2 from your IaC.
  3. Point the ALB or NLB health checks at a dependency-aware path, and confirm each Region can absorb the other's peak.
  4. Enable client IP preservation where supported and update security groups for real client addresses.
  5. Add application heartbeats for any connection that can sit idle longer than 340 seconds.
  6. Rehearse a drain with traffic dials and an unplanned failover, measuring how long new connections take to move.
Key takeaway: Global Accelerator puts two static anycast IPs in front of your Regions, terminates TCP at the nearest edge and carries traffic over the AWS backbone. Failover happens behind the IP: unhealthy endpoints stop getting new connections without any DNS change. Traffic dials limit a Region's share of its own traffic, weights split traffic inside a Region, established flows persist until they close or hit the 340-second TCP idle timeout, and with no healthy endpoints it fails open. Use it when clients need fixed IPs or faster-than-DNS failover, and size every Region to take over.