Google Cloud's global external load balancers give one anycast IP address to users worldwide and send each request to a healthy backend, usually in the nearest region with spare capacity. The family covers the global external Application Load Balancer and the global external proxy Network Load Balancer, both using the EXTERNAL_MANAGED scheme, plus their older classic modes, which use EXTERNAL. All of them run on Google Front Ends (GFEs), and the global modes require the Premium network tier.
The resource chain, the product families and how to choose between them are covered in GCP load balancing. This article goes one level down, into what happens between the client and your backend: how requests cross the two layers of GFEs, how the load balancer decides which region and zone gets traffic and when it spills over, the knobs that control that decision, advanced routing for canaries, the timeouts that cause intermittent errors, and how to migrate off the classic mode.
The request path: anycast and two layers of GFEs
The load balancer's IP is announced from many Google edge locations at once, so the internet routes each client to a nearby point of presence; anycast explains how that works. There a first-layer GFE terminates the client's TCP connection and TLS session. The client's handshakes therefore cross only the short distance to the edge, and the long haul to the backend region runs over Google's backbone on connections the GFEs keep open.
The first-layer GFE forwards the request to a second-layer GFE in a region it selects from the regions that have backends, preferring the lowest round-trip time and moving to others when the preferred region is at capacity. The second-layer GFE chooses a zone and then an endpoint, using the capacity and health it tracks for each backend. The global external Application Load Balancer uses the open-source Envoy proxy for its advanced traffic management, which is why it supports richer routing than the classic mode.
Two consequences follow. Your backends see connections from Google, not from clients, and must read the client address from X-Forwarded-For. And every setting has a layer: certificates and URL-map routing apply at the edge, while capacity, health and endpoint selection belong to the backend service.
Choosing a region: service load balancing policies
Region and zone selection is governed by a service load balancing policy attached to the backend service. If you attach none, the load balancer behaves as WATERFALL_BY_REGION. In all three algorithms traffic goes to the region closest to the user first and overflows to other regions at capacity; they differ in how the proxies spread requests across the zones of a region:
| Algorithm | Behaviour within the region | Good for |
|---|---|---|
WATERFALL_BY_REGION (default) | The proxies in the closest region fill backends in proportion to their configured target capacities | A balanced default |
SPRAY_TO_REGION | Each proxy sends to all endpoints in all zones of the region, with no preference for a shorter round trip | Avoiding hot zones, warm caches in every zone |
WATERFALL_BY_ZONE | Each proxy strongly prefers endpoints in the closest possible zone | Minimising cross-zone traffic when capacity per zone is ample |
Waterfall routing means the capacity you declare is a steering control. Worked example: two regions, each a managed instance group with 10 instances and a RATE capacity of 300 requests per second per instance, so 3,000 per region. At the evening peak European users generate 4,500 requests per second. The load balancer sends 3,000 to europe-west1 and spills the remaining 1,500 to us-east1, adding a transatlantic round trip to one request in three. If the declared rate were 1,000 per instance, it would never spill, and the European instances would drown while the US ones idled. Set the rate from load tests and let autoscaling grow the groups, because spillover is a safety valve, not a plan.
Health-driven draining, failover and preferred backends
Waterfall routing reacts to capacity. Health needs separate controls, all configured in the same policy:
- Auto-capacity draining. When fewer than 25% of a backend's instances or endpoints pass health checks, its capacity is set to zero and traffic moves elsewhere. It returns to its normal capacity when 35% or more pass for at least 60 seconds. As a safety limit, the documentation says draining applies only while the backends eligible to be drained are fewer than 50% of all backends, so a global health-check mistake cannot drain everything.
- Failover threshold.
failoverHealthThresholdsets the percentage of healthy endpoints below which a backend is treated as failing over, from 1 to 99, default 70. - Preferred backends. Marking a backend
PREFERREDmakes the load balancer use it fully before spilling to others, for example to use committed capacity in one region first. - Isolation. A Preview feature lets you set isolation to
STRICTwithREGIONgranularity to stop cross-region overflow, for data-residency or cost reasons. Being Preview, do not rely on it for compliance without checking its current status.
# service-lb-policy.yaml : attach to a global backend service
name: projects/PROJECT_ID/locations/global/serviceLbPolicies/web-policy
loadBalancingAlgorithm: WATERFALL_BY_REGION # default; or SPRAY_TO_REGION, WATERFALL_BY_ZONE
autoCapacityDrain:
enable: True # drain a backend when fewer than 25% of its endpoints are healthy
failoverConfig:
failoverHealthThreshold: 70 # default; percentage of healthy endpoints before failovergcloud network-services service-lb-policies import web-policy \
--source=service-lb-policy.yaml --location=global
gcloud compute backend-services update web-bs --global \
--service-lb-policy=projects/PROJECT_ID/locations/global/serviceLbPolicies/web-policy
# Fill the regional backend you want used first before spilling elsewhere
gcloud compute backend-services update-backend web-bs --global \
--instance-group=web-eu --instance-group-region=europe-west1 --preference=PREFERRED
Inside a region: endpoint choice and outlier detection
After the zone is chosen, the backend service's localityLbPolicy picks the endpoint. Values in the API include ROUND_ROBIN (the default), LEAST_REQUEST, which compares two random hosts and picks the one with fewer active requests, and consistent-hash policies such as RING_HASH and MAGLEV for affinity. Not every value applies to every load balancer mode, so check the backend service API reference for yours. LEAST_REQUEST is the better default when request costs vary widely.
Health checks catch dead endpoints; outlier detection catches ones that answer health checks but fail real requests. The load balancer ejects an endpoint after consecutive errors for a base ejection time of 30 seconds by default, multiplied by the number of times it has been ejected, and never ejects more than 50% of endpoints by default. The cap matters: if half your fleet fails at once, the problem is systemic and ejection would only concentrate load on the rest.
Advanced routing: canaries, header routes and retries
The global external Application Load Balancer's URL map supports route rules with priorities, header and query matching, weighted splits between backend services, retries, timeouts, URL rewrites, fault injection and request mirroring. Weights range from 0 to 1000 per backend service, so the precision of a split is fine enough for 0.5% canaries.
# url-map.yaml (excerpt): canary, tester routing, retries and a timeout
defaultService: projects/PROJECT_ID/global/backendServices/web-v1
hostRules:
- hosts: ["www.example.com"]
pathMatcher: web
pathMatchers:
- name: web
defaultService: projects/PROJECT_ID/global/backendServices/web-v1
routeRules:
- priority: 1 # internal testers always get v2
matchRules:
- prefixMatch: /
headerMatches:
- headerName: x-canary
exactMatch: "always"
service: projects/PROJECT_ID/global/backendServices/web-v2
- priority: 2 # everyone else: 5% to v2
matchRules:
- prefixMatch: /
routeAction:
weightedBackendServices:
- backendService: projects/PROJECT_ID/global/backendServices/web-v1
weight: 950
- backendService: projects/PROJECT_ID/global/backendServices/web-v2
weight: 50
retryPolicy:
retryConditions: ["connect-failure", "gateway-error"]
numRetries: 2
perTryTimeout: {seconds: 5}
timeout: {seconds: 20}The rules are evaluated by priority. Internal testers who send the x-canary header always reach v2; other traffic is split 95/5. Retries apply only to conditions you list. A connect-failure means the request never reached the backend, so it is safe to retry anywhere. gateway-error covers 502, 503 and 504 responses, and a 504 or an application's own 503 can arrive after the backend did the work, so allow it only on routes whose requests are idempotent, as this read-mostly page route is. Retrying all 5xx responses on a write route can repeat payments or orders. Apply the file with gcloud compute url-maps import and roll forward by editing the weights, following the practice in canary deployments. Request mirroring copies traffic to a shadow service for testing, but it is not supported for internet NEGs, serverless NEGs or Private Service Connect backends.
Timeouts and keepalives: the source of mystery 502s
| Setting | Value (global external ALB) | What to do |
|---|---|---|
| Client HTTP keepalive | 610 s default, configurable 5 to 1,200 s | Rarely change; lower only to shed idle clients |
| Backend HTTP keepalive | 600 s, fixed | Make backend servers keep idle connections longer than 600 s |
| Backend service timeout | 30 s default (serverless NEGs differ); up to 2,147,483,647 s | Raise for long requests, streaming or WebSockets |
The fixed 600-second backend keepalive is the key number. The load balancer reuses idle connections to your servers for up to ten minutes. Many web servers close idle connections after a few seconds. When the server closes a connection just as the load balancer sends a new request on it, the request fails and the client sees an error that looks random and correlates with low traffic. Configure your servers to idle longer than the load balancer does:
# nginx on the backends: idle longer than the load balancer's fixed 600 s backend keepalive
keepalive_timeout 620s;
# gunicorn equivalent
# gunicorn app:app --keep-alive 620The backend service timeout is measured from the first byte of the request sent to the backend until the last byte of the response returns, so it caps the whole response, not idle time. Streaming endpoints need a timeout sized to the longest stream. When the load balancer closes a connection, it does so cleanly with a FIN or an HTTP/2 GOAWAY frame rather than interrupting active requests.
Security and caching at the edge
Because the first layer terminates traffic at the edge, attacks and cacheable traffic can be stopped there. Cloud Armor security policies attach to backend services and filter requests before they reach a region, as covered in Cloud Armor. Cloud CDN, enabled on a backend service or backend bucket, serves cached responses from the edge, as described in Cloud CDN. Both work best with the routing above: route static paths to a CDN-enabled backend bucket, and apply stricter Armor rules to API paths.
Migrating from classic
Classic Application Load Balancers use the EXTERNAL scheme and lack much of the advanced traffic management above. Google supports migrating them in place, keeping the IP address, by moving backend services and backend buckets first and forwarding rules last. Each resource steps through --external-managed-migration-state values: PREPARE, optionally TEST_BY_PERCENTAGE with --external-managed-migration-testing-percentage, then TEST_ALL_TRAFFIC, and finally the scheme changes to EXTERNAL_MANAGED. Google advises allowing about six minutes between state changes, and rollback remains possible for 90 days after the scheme change.
Check for blockers first: the global mode does not support the Standard network tier, so a classic load balancer on Standard tier cannot migrate as it is. Then test for the documented behaviour differences, which break clients silently. The global mode returns HTTP 503 rather than 502 when all backends are unhealthy, so update alerts keyed on 502. It lowercases all header keys, whereas classic preserved their case, which breaks code that compares header names case-sensitively. And the percentage test phase splits traffic between two infrastructures, which breaks client-IP session affinity while it lasts.
Failure modes
| Symptom | Likely cause | Fix |
|---|---|---|
| Sporadic 502s at low traffic | Backend closes idle connections before 600 s | Backend keepalive above 600 s |
| Requests cut at 30 s | Default backend service timeout | Raise the timeout on that backend service |
| Nearest region overloaded, others idle | Declared capacity above real capacity | Set RATE from load tests; autoscale |
| Traffic sprays to a far region | Declared capacity too low, or auto-capacity drain triggered | Check capacity and health per backend |
| Retries duplicate writes | Retry on 5xx or gateway-error for non-idempotent routes | Retry only connect-failure on write routes |
| Alerts silent after migration | 503 replaced 502; header case changed | Update alerts and header handling |
Trade-offs
One global anycast IP with edge termination gives the lowest handshake latency, a single certificate and DNS entry, and automatic cross-region failover. The cost is Premium tier pricing, backends that never see client IPs directly, and a traffic distribution that depends on capacity numbers you must keep accurate. Waterfall routing favours latency; spraying favours even load. Aggressive draining protects users from failing backends but can push a whole region's load onto another at once, so size regions to absorb a neighbour's overflow.
What to do next
- Confirm each load balancer's scheme; plan migration for any still on
EXTERNAL, starting with a 1% test percentage. - Set every backend's declared capacity from load tests and verify overflow in a game day by shrinking one region.
- Attach a service load balancing policy with auto-capacity draining, and decide explicitly between waterfall and spray.
- Set backend server keepalive above 600 seconds and review backend service timeouts for streaming or long requests.
- Move canaries into URL-map weighted splits with a header override for testers; retry connect-failure everywhere and gateway-error only on idempotent routes.
- Enable outlier detection and load balancer logging, and alert on both 502 and 503 responses.