What Google Cloud used to call the global HTTP(S) Load Balancer is now the global external Application Load Balancer. It terminates HTTP and HTTPS at Google's edge on a single anycast IP address and proxies requests to backends in any region. There are two generations under that name: the current one, selected with the load-balancing scheme EXTERNAL_MANAGED, and the classic one, scheme EXTERNAL. Everything on this page refers to EXTERNAL_MANAGED unless it says otherwise; some defaults differ for classic, so check its documentation rather than assuming.
Choosing between load balancer types is covered in GCP load balancing, and the edge architecture (Google Front Ends, cross-region failover, advanced routing) in the global external load balancer family. This page is the operating manual: how to declare one as code, how certificates actually get issued, how to test routing before you ship it, and how to read the logs when users start seeing 502s.
The object chain
A Google Cloud load balancer is not one resource but a chain of small ones, and nearly every outage maps to one link in it. A global forwarding rule binds an IP address and port to a target proxy. The target HTTPS proxy holds the certificates and SSL policy and points to a URL map, which routes each request by host, path, headers or query to a backend service (or a backend bucket for Cloud Storage). The backend service carries the timeout, logging, session affinity, Cloud Armor policy and CDN settings, and lists its backends: managed instance groups, zonal network endpoint groups (NEGs) for GKE, serverless NEGs for Cloud Run, or internet NEGs. Instance-group and zonal-NEG backends have a health check, and the VPC needs a firewall rule that lets Google's proxies and probers in.
Declaring it in Terraform
Because the chain is many objects, declare it as code. The core of a two-host setup, with an API on a managed instance group and a web front end on Cloud Run:
resource "google_compute_global_address" "web" {
name = "web-ip"
}
resource "google_compute_managed_ssl_certificate" "web" {
name = "web-cert"
managed {
domains = ["shop.example.com", "api.example.com"]
}
}
resource "google_compute_health_check" "api" {
name = "api-hc"
http_health_check {
port = 8080
request_path = "/healthz"
}
}
resource "google_compute_backend_service" "api" {
name = "api-backend"
load_balancing_scheme = "EXTERNAL_MANAGED"
protocol = "HTTP"
port_name = "http"
timeout_sec = 30
health_checks = [google_compute_health_check.api.id]
backend {
group = google_compute_region_instance_group_manager.api.instance_group
balancing_mode = "UTILIZATION"
capacity_scaler = 1.0
}
log_config {
enable = true
sample_rate = 1.0
}
}
resource "google_compute_url_map" "web" {
name = "web-map"
default_service = google_compute_backend_service.web.id # serverless NEG backend
host_rule {
hosts = ["api.example.com"]
path_matcher = "api"
}
path_matcher {
name = "api"
default_service = google_compute_backend_service.api.id
}
test {
host = "api.example.com"
path = "/v1/orders"
service = google_compute_backend_service.api.id
}
}
resource "google_compute_target_https_proxy" "web" {
name = "web-https"
url_map = google_compute_url_map.web.id
ssl_certificates = [google_compute_managed_ssl_certificate.web.id]
}
resource "google_compute_global_forwarding_rule" "https" {
name = "web-https"
load_balancing_scheme = "EXTERNAL_MANAGED"
ip_address = google_compute_global_address.web.id
port_range = "443"
target = google_compute_target_https_proxy.web.id
}
resource "google_compute_firewall" "lb_to_api" {
name = "allow-lb-and-probes"
network = google_compute_network.main.id
direction = "INGRESS"
source_ranges = ["35.191.0.0/16", "130.211.0.0/22"]
target_tags = ["api"]
allow {
protocol = "tcp"
ports = ["8080"]
}
}The network, the instance group manager and the Cloud Run backend service web are declared elsewhere. Port 80 gets its own URL map whose only job is a default_url_redirect block with https_redirect = true and strip_query = false, plus a target HTTP proxy and a second forwarding rule on the same address. Keep the redirect map separate from the serving map so a routing change can never accidentally serve plain HTTP. Note the test block: Terraform sends it with the URL map, and Google Cloud rejects the update if a test fails.
Certificates: how issuance actually works
Google-managed certificates are free and renew themselves, but issuance has preconditions that catch almost everyone once. For a certificate attached directly to the proxy, each domain must already resolve in public DNS to the load balancer's IP, the forwarding rule must listen on 443, and the certificate must be attached to the target proxy. Each domain reports its own status: PROVISIONING, ACTIVE, FAILED_NOT_VISIBLE, FAILED_CAA_CHECKING, FAILED_CAA_FORBIDDEN or FAILED_RATE_LIMITED. Read them with gcloud compute ssl-certificates describe web-cert --global.
FAILED_NOT_VISIBLE almost always means DNS does not point at the load balancer yet, or the certificate is not attached. The CAA statuses mean a CAA record on your domain does not authorise the certificate authorities Google uses; the documentation lists the values to add. Google's documentation says provisioning can take up to 60 minutes after DNS and load balancer changes have propagated, and an active certificate may take about 30 more minutes to be served everywhere.
That DNS precondition creates a chicken-and-egg problem when migrating a live domain: you cannot get the certificate until DNS points at the new load balancer, and you do not want to point DNS there until it has a certificate. Certificate Manager solves it with DNS authorization: you add a CNAME record that proves control of the domain, the certificate is issued while traffic still goes to the old system, and you attach it to the proxy through a certificate map before cutting DNS over. Use that path for any migration where downtime matters.
URL maps with tests
URL maps grow host rules, path matchers, header matches, rewrites and redirects, and a wrong rule sends traffic to the wrong service with no error at all. Test them like code. A URL map can carry a tests list; each test names a host and path and either the backend service it should reach or the expected redirect:
# web-map.yaml (export with: gcloud compute url-maps export web-map --global)
tests:
- host: api.example.com
path: /v1/orders
service: projects/my-proj/global/backendServices/api-backend
- host: shop.example.com
path: /
service: projects/my-proj/global/backendServices/web-backend
- host: example.com
path: /pricing
expectedOutputUrl: https://shop.example.com/pricing
expectedRedirectResponseCode: 301gcloud compute url-maps validate --source=web-map.yaml \
--load-balancing-scheme=EXTERNAL_MANAGED --globalvalidate runs static checks and the tests without creating or changing the URL map, so it belongs in CI on every routing pull request. Add a test for every rule you add, and for every bug a routing rule has ever caused.
Health checks, firewalls and backends
Health-check probes for this load balancer come from 35.191.0.0/16 for IPv4 (and 2600:2d00:1:b029::/64 for IPv6). The documentation also lists 130.211.0.0/22 alongside 35.191.0.0/16 as sources of proxied requests to instance group and zonal NEG backends, so the firewall rule above allows both, on the serving port and the health-check port. Without it, the VPC's implied deny drops the probes, every backend is marked unhealthy and the load balancer has nothing to send traffic to.
Point the health check at an endpoint that reflects whether the instance can serve, not merely whether the process is up, but keep it cheap and free of downstream dependencies; otherwise a database blip marks every backend unhealthy at once. Serverless NEGs for Cloud Run have no health check: Cloud Run manages instance health itself, as covered in Cloud Run in depth. To drain an instance group for maintenance, lower its capacity_scaler toward zero rather than deleting it, so in-flight requests finish.
Timeouts and retries
| Setting (EXTERNAL_MANAGED) | Value | Operational consequence |
|---|---|---|
| Backend service timeout | 30 s default for most backend types; configurable | Long requests and idle WebSockets are cut at this value |
| Client HTTP keepalive | 610 s default, configurable 5 to 1,200 s | Rarely needs changing |
| Backend HTTP keepalive | 600 s, fixed | Backend servers must keep idle connections open longer than 600 s |
| WebSocket lifetime | Active connections closed after 24 hours | Clients must reconnect |
| Default retries | Requests without a body (for example GET) that get 502, 503 or 504 are retried once | POST is not retried; make retries safe with idempotency |
The fixed 600-second backend keepalive is the most important number on the page. Many web servers default to idle timeouts of a few seconds to a minute. When the backend closes an idle connection that the proxy is about to reuse, the request on it fails, and the log says backend_connection_closed_before_data_sent_to_client. Set the server's keepalive or idle timeout above 600 seconds, for example 620.
Reading the logs
With logging enabled on the backend service, every request becomes a log entry with resource type http_load_balancer. The standard httpRequest fields hold status, latency, URL and user agent. The field that matters most is jsonPayload.statusDetails, which says why the load balancer returned what it returned. Start every investigation by grouping errors on it in Logs Explorer:
resource.type="http_load_balancer"
httpRequest.status>=500
jsonPayload.statusDetails!="response_sent_by_backend"Two Cloud Monitoring metrics separate client-side from backend-side latency. loadbalancing.googleapis.com/https/backend_latencies measures from the last byte of the request sent to the backend to the last byte of the response received. loadbalancing.googleapis.com/https/total_latencies measures the whole exchange with the client. If both rise, the backend is slow. If only total latency rises, look at the client path: large responses, slow clients or uploads. The default log sample rate is 1.0, every request; at high traffic lower it for cost, keeping enough samples to see rare errors. See Cloud Logging for sinks and retention.
A statusDetails runbook
| statusDetails | What it means | First things to check |
|---|---|---|
response_sent_by_backend | The backend produced the response, including any 5xx | Application logs; this is not a load balancer problem |
failed_to_pick_backend | No healthy backend was available | Health-check status, firewall for 35.191.0.0/16, capacity_scaler, empty groups |
failed_to_connect_to_backend | The proxy could not open a connection | Firewall, named port and port_name, process listening, connection limits |
backend_timeout | The backend did not respond within the backend service timeout | Slow endpoints, timeout_sec, long polls and WebSockets |
backend_connection_closed_before_data_sent_to_client | The backend closed the connection before responding | Server keepalive under 600 s, crashes, restarts during deploys |
client_disconnected_before_any_response | The client gave up first | Total latency, client timeouts, mobile networks |
backend_early_response_with_non_error_status | The backend answered before reading the whole request body | Upload handlers that respond early |
Worked example: 502s during every deploy
A team sees a burst of 502s for about a minute every time they roll out a new version of the API instance group. Grouping on statusDetails shows two values. Most errors are backend_connection_closed_before_data_sent_to_client, the rest failed_to_connect_to_backend.
The first value points at connections closed under the proxy. The server was being stopped with SIGKILL by the rollout, dropping connections the proxy had pooled, and it also had a 75-second idle timeout, so occasional errors showed up even between deploys. The second value appears because instances were removed before health checks had marked them unhealthy, so the proxy kept trying them. The fix had three parts: raise the server's idle timeout to 620 seconds; handle SIGTERM by failing the health check first, waiting for traffic to drain, then closing; and configure connection draining on the backend service so removed instances get time to finish. The next deploy produced no load-balancer-originated 502s. Because GETs are retried once on 502, users of read endpoints barely noticed the old problem, which is why it had gone unfixed: the POST requests were the ones failing.
Trade-offs and gotchas
- Header case. The global load balancer may convert header names to lowercase; code that matches header names case-sensitively breaks during migration from classic.
- Client IP. Backends see the proxy's address; the client IP is in
X-Forwarded-For, which ends with the client IP and the load balancer IP. Parse from the right; earlier values can be spoofed. - Edge security and caching attach to backend services: Cloud Armor policies, covered in Cloud Armor, and Cloud CDN.
- Cost. You pay for forwarding rules and data processed. For a single regional service with no global users, a regional load balancer can be simpler; for multi-region failover and one IP worldwide, the global one is the right tool.
What to do next
- Confirm your load balancer's scheme is
EXTERNAL_MANAGED; if it is classic, plan the migration and test header handling. - Put the whole object chain in Terraform, with the port-80 redirect on its own URL map.
- Use Certificate Manager with DNS authorization for any domain you migrate, and alert on any domain status that is not
ACTIVE. - Add URL-map tests for every routing rule and run
gcloud compute url-maps validatein CI. - Check the firewall allows
35.191.0.0/16and130.211.0.0/22on serving and health-check ports, and that the health endpoint has no downstream dependencies. - Set every backend server's keepalive above 600 seconds and make shutdown fail health checks before closing connections.
- Enable logging, save the
statusDetailsquery, and chart backend versus total latency on one dashboard.