Google Cloud's global external load balancers give one anycast IP address to users worldwide and send each request to a healthy backend, usually in the nearest region with spare capacity. The family covers the global external Application Load Balancer and the global external proxy Network Load Balancer, both using the EXTERNAL_MANAGED scheme, plus their older classic modes, which use EXTERNAL. All of them run on Google Front Ends (GFEs), and the global modes require the Premium network tier.

The resource chain, the product families and how to choose between them are covered in GCP load balancing. This article goes one level down, into what happens between the client and your backend: how requests cross the two layers of GFEs, how the load balancer decides which region and zone gets traffic and when it spills over, the knobs that control that decision, advanced routing for canaries, the timeouts that cause intermittent errors, and how to migrate off the classic mode.

Advertisement

The request path: anycast and two layers of GFEs

Client in Parisresolves one IPClient in Ohiosame anycast IPanycastanycastFirst-layer GFEedge PoP: TCP + TLS endFirst-layer GFEedge PoP: TCP + TLS endnearestnearestspilloverSecond-layer GFE, europe-west1picks zone and endpointSecond-layer GFE, us-east1picks zone and endpointEU backendsMIGs or NEGs, 2 zonesUS backendsMIGs or NEGs, 2 zonesAttached at the edgeURL map routing, Cloud Armor, Cloud CDN, certificatesservice LB policy: WATERFALL_BY_REGION by default
One anycast IP is announced from Google's edge. First-layer GFEs terminate TCP and TLS near the client and forward to second-layer GFEs in a region chosen by round-trip time and capacity; second-layer GFEs pick the zone and endpoint. Routing, security and caching policies apply at the edge.

The load balancer's IP is announced from many Google edge locations at once, so the internet routes each client to a nearby point of presence; anycast explains how that works. There a first-layer GFE terminates the client's TCP connection and TLS session. The client's handshakes therefore cross only the short distance to the edge, and the long haul to the backend region runs over Google's backbone on connections the GFEs keep open.

The first-layer GFE forwards the request to a second-layer GFE in a region it selects from the regions that have backends, preferring the lowest round-trip time and moving to others when the preferred region is at capacity. The second-layer GFE chooses a zone and then an endpoint, using the capacity and health it tracks for each backend. The global external Application Load Balancer uses the open-source Envoy proxy for its advanced traffic management, which is why it supports richer routing than the classic mode.

Two consequences follow. Your backends see connections from Google, not from clients, and must read the client address from X-Forwarded-For. And every setting has a layer: certificates and URL-map routing apply at the edge, while capacity, health and endpoint selection belong to the backend service.

Choosing a region: service load balancing policies

Region and zone selection is governed by a service load balancing policy attached to the backend service. If you attach none, the load balancer behaves as WATERFALL_BY_REGION. In all three algorithms traffic goes to the region closest to the user first and overflows to other regions at capacity; they differ in how the proxies spread requests across the zones of a region:

AlgorithmBehaviour within the regionGood for
WATERFALL_BY_REGION (default)The proxies in the closest region fill backends in proportion to their configured target capacitiesA balanced default
SPRAY_TO_REGIONEach proxy sends to all endpoints in all zones of the region, with no preference for a shorter round tripAvoiding hot zones, warm caches in every zone
WATERFALL_BY_ZONEEach proxy strongly prefers endpoints in the closest possible zoneMinimising cross-zone traffic when capacity per zone is ample

Waterfall routing means the capacity you declare is a steering control. Worked example: two regions, each a managed instance group with 10 instances and a RATE capacity of 300 requests per second per instance, so 3,000 per region. At the evening peak European users generate 4,500 requests per second. The load balancer sends 3,000 to europe-west1 and spills the remaining 1,500 to us-east1, adding a transatlantic round trip to one request in three. If the declared rate were 1,000 per instance, it would never spill, and the European instances would drown while the US ones idled. Set the rate from load tests and let autoscaling grow the groups, because spillover is a safety valve, not a plan.

Advertisement

Health-driven draining, failover and preferred backends

Waterfall routing reacts to capacity. Health needs separate controls, all configured in the same policy:

  • Auto-capacity draining. When fewer than 25% of a backend's instances or endpoints pass health checks, its capacity is set to zero and traffic moves elsewhere. It returns to its normal capacity when 35% or more pass for at least 60 seconds. As a safety limit, the documentation says draining applies only while the backends eligible to be drained are fewer than 50% of all backends, so a global health-check mistake cannot drain everything.
  • Failover threshold. failoverHealthThreshold sets the percentage of healthy endpoints below which a backend is treated as failing over, from 1 to 99, default 70.
  • Preferred backends. Marking a backend PREFERRED makes the load balancer use it fully before spilling to others, for example to use committed capacity in one region first.
  • Isolation. A Preview feature lets you set isolation to STRICT with REGION granularity to stop cross-region overflow, for data-residency or cost reasons. Being Preview, do not rely on it for compliance without checking its current status.
# service-lb-policy.yaml : attach to a global backend service
name: projects/PROJECT_ID/locations/global/serviceLbPolicies/web-policy
loadBalancingAlgorithm: WATERFALL_BY_REGION   # default; or SPRAY_TO_REGION, WATERFALL_BY_ZONE
autoCapacityDrain:
  enable: True            # drain a backend when fewer than 25% of its endpoints are healthy
failoverConfig:
  failoverHealthThreshold: 70   # default; percentage of healthy endpoints before failover
gcloud network-services service-lb-policies import web-policy \
    --source=service-lb-policy.yaml --location=global

gcloud compute backend-services update web-bs --global \
    --service-lb-policy=projects/PROJECT_ID/locations/global/serviceLbPolicies/web-policy

# Fill the regional backend you want used first before spilling elsewhere
gcloud compute backend-services update-backend web-bs --global \
    --instance-group=web-eu --instance-group-region=europe-west1 --preference=PREFERRED

Inside a region: endpoint choice and outlier detection

After the zone is chosen, the backend service's localityLbPolicy picks the endpoint. Values in the API include ROUND_ROBIN (the default), LEAST_REQUEST, which compares two random hosts and picks the one with fewer active requests, and consistent-hash policies such as RING_HASH and MAGLEV for affinity. Not every value applies to every load balancer mode, so check the backend service API reference for yours. LEAST_REQUEST is the better default when request costs vary widely.

Health checks catch dead endpoints; outlier detection catches ones that answer health checks but fail real requests. The load balancer ejects an endpoint after consecutive errors for a base ejection time of 30 seconds by default, multiplied by the number of times it has been ejected, and never ejects more than 50% of endpoints by default. The cap matters: if half your fleet fails at once, the problem is systemic and ejection would only concentrate load on the rest.

Advanced routing: canaries, header routes and retries

The global external Application Load Balancer's URL map supports route rules with priorities, header and query matching, weighted splits between backend services, retries, timeouts, URL rewrites, fault injection and request mirroring. Weights range from 0 to 1000 per backend service, so the precision of a split is fine enough for 0.5% canaries.

# url-map.yaml (excerpt): canary, tester routing, retries and a timeout
defaultService: projects/PROJECT_ID/global/backendServices/web-v1
hostRules:
- hosts: ["www.example.com"]
  pathMatcher: web
pathMatchers:
- name: web
  defaultService: projects/PROJECT_ID/global/backendServices/web-v1
  routeRules:
  - priority: 1                                  # internal testers always get v2
    matchRules:
    - prefixMatch: /
      headerMatches:
      - headerName: x-canary
        exactMatch: "always"
    service: projects/PROJECT_ID/global/backendServices/web-v2
  - priority: 2                                  # everyone else: 5% to v2
    matchRules:
    - prefixMatch: /
    routeAction:
      weightedBackendServices:
      - backendService: projects/PROJECT_ID/global/backendServices/web-v1
        weight: 950
      - backendService: projects/PROJECT_ID/global/backendServices/web-v2
        weight: 50
      retryPolicy:
        retryConditions: ["connect-failure", "gateway-error"]
        numRetries: 2
        perTryTimeout: {seconds: 5}
      timeout: {seconds: 20}

The rules are evaluated by priority. Internal testers who send the x-canary header always reach v2; other traffic is split 95/5. Retries apply only to conditions you list. A connect-failure means the request never reached the backend, so it is safe to retry anywhere. gateway-error covers 502, 503 and 504 responses, and a 504 or an application's own 503 can arrive after the backend did the work, so allow it only on routes whose requests are idempotent, as this read-mostly page route is. Retrying all 5xx responses on a write route can repeat payments or orders. Apply the file with gcloud compute url-maps import and roll forward by editing the weights, following the practice in canary deployments. Request mirroring copies traffic to a shadow service for testing, but it is not supported for internet NEGs, serverless NEGs or Private Service Connect backends.

Timeouts and keepalives: the source of mystery 502s

SettingValue (global external ALB)What to do
Client HTTP keepalive610 s default, configurable 5 to 1,200 sRarely change; lower only to shed idle clients
Backend HTTP keepalive600 s, fixedMake backend servers keep idle connections longer than 600 s
Backend service timeout30 s default (serverless NEGs differ); up to 2,147,483,647 sRaise for long requests, streaming or WebSockets

The fixed 600-second backend keepalive is the key number. The load balancer reuses idle connections to your servers for up to ten minutes. Many web servers close idle connections after a few seconds. When the server closes a connection just as the load balancer sends a new request on it, the request fails and the client sees an error that looks random and correlates with low traffic. Configure your servers to idle longer than the load balancer does:

# nginx on the backends: idle longer than the load balancer's fixed 600 s backend keepalive
keepalive_timeout 620s;

# gunicorn equivalent
#   gunicorn app:app --keep-alive 620

The backend service timeout is measured from the first byte of the request sent to the backend until the last byte of the response returns, so it caps the whole response, not idle time. Streaming endpoints need a timeout sized to the longest stream. When the load balancer closes a connection, it does so cleanly with a FIN or an HTTP/2 GOAWAY frame rather than interrupting active requests.

Security and caching at the edge

Because the first layer terminates traffic at the edge, attacks and cacheable traffic can be stopped there. Cloud Armor security policies attach to backend services and filter requests before they reach a region, as covered in Cloud Armor. Cloud CDN, enabled on a backend service or backend bucket, serves cached responses from the edge, as described in Cloud CDN. Both work best with the routing above: route static paths to a CDN-enabled backend bucket, and apply stricter Armor rules to API paths.

Migrating from classic

Classic Application Load Balancers use the EXTERNAL scheme and lack much of the advanced traffic management above. Google supports migrating them in place, keeping the IP address, by moving backend services and backend buckets first and forwarding rules last. Each resource steps through --external-managed-migration-state values: PREPARE, optionally TEST_BY_PERCENTAGE with --external-managed-migration-testing-percentage, then TEST_ALL_TRAFFIC, and finally the scheme changes to EXTERNAL_MANAGED. Google advises allowing about six minutes between state changes, and rollback remains possible for 90 days after the scheme change.

Check for blockers first: the global mode does not support the Standard network tier, so a classic load balancer on Standard tier cannot migrate as it is. Then test for the documented behaviour differences, which break clients silently. The global mode returns HTTP 503 rather than 502 when all backends are unhealthy, so update alerts keyed on 502. It lowercases all header keys, whereas classic preserved their case, which breaks code that compares header names case-sensitively. And the percentage test phase splits traffic between two infrastructures, which breaks client-IP session affinity while it lasts.

Failure modes

SymptomLikely causeFix
Sporadic 502s at low trafficBackend closes idle connections before 600 sBackend keepalive above 600 s
Requests cut at 30 sDefault backend service timeoutRaise the timeout on that backend service
Nearest region overloaded, others idleDeclared capacity above real capacitySet RATE from load tests; autoscale
Traffic sprays to a far regionDeclared capacity too low, or auto-capacity drain triggeredCheck capacity and health per backend
Retries duplicate writesRetry on 5xx or gateway-error for non-idempotent routesRetry only connect-failure on write routes
Alerts silent after migration503 replaced 502; header case changedUpdate alerts and header handling

Trade-offs

One global anycast IP with edge termination gives the lowest handshake latency, a single certificate and DNS entry, and automatic cross-region failover. The cost is Premium tier pricing, backends that never see client IPs directly, and a traffic distribution that depends on capacity numbers you must keep accurate. Waterfall routing favours latency; spraying favours even load. Aggressive draining protects users from failing backends but can push a whole region's load onto another at once, so size regions to absorb a neighbour's overflow.

What to do next

  1. Confirm each load balancer's scheme; plan migration for any still on EXTERNAL, starting with a 1% test percentage.
  2. Set every backend's declared capacity from load tests and verify overflow in a game day by shrinking one region.
  3. Attach a service load balancing policy with auto-capacity draining, and decide explicitly between waterfall and spray.
  4. Set backend server keepalive above 600 seconds and review backend service timeouts for streaming or long requests.
  5. Move canaries into URL-map weighted splits with a header override for testers; retry connect-failure everywhere and gateway-error only on idempotent routes.
  6. Enable outlier detection and load balancer logging, and alert on both 502 and 503 responses.
Key takeaway: Google's global external load balancers terminate connections at first-layer GFEs near the client, then hand requests to second-layer GFEs in the nearest region with capacity, which pick a zone and endpoint. Declared capacity steers that waterfall, service load balancing policies add auto-capacity draining, failover thresholds and preferred backends, and URL maps add weighted canaries and retries. Keep backend keepalives above the fixed 600 seconds, size timeouts deliberately, and migrate classic load balancers with the staged states while testing for the 503 and header-case changes.