Google Cloud does not have one load balancer; it has a family of them that share a resource model but differ in where traffic is terminated, which protocols they carry, whether they are global or regional, and whether they face the internet or only your VPC. Picking the wrong member is expensive to undo, because the choice decides whether your backends see client IP addresses, whether you can route on URL paths, whether you need a special subnet, and which features such as Cloud Armor or Cloud CDN are available.

This article builds the model from the resources up, maps the product family, gives a decision procedure, walks through a complete global HTTPS setup, and covers how traffic is distributed, what breaks in production and how to operate it. General load-balancing concepts such as L4 versus L7 and health checking are covered in cloud load balancer architecture; this page is about how Google Cloud implements them.

Advertisement

The resource model: a chain of small objects

Every GCP load balancer is assembled from separate resources rather than created as one object. Understanding the chain makes both the console and the error messages readable.

  • Forwarding rule: the frontend. An IP address, protocol, port range and a load-balancing scheme that fixes which product this is.
  • Target proxy (proxy load balancers only): terminates the client connection. An HTTPS proxy holds certificates; a TCP or SSL proxy handles non-HTTP traffic.
  • URL map (Application Load Balancers only): routing rules by host and path, plus features such as redirects, rewrites and weighted traffic splitting.
  • Backend service: a group of backends plus the policy for them: protocol, health check, balancing mode, session affinity, timeouts, logging, and attachments such as a Cloud Armor policy or Cloud CDN.
  • Backends: managed or unmanaged instance groups, or network endpoint groups (NEGs) that point at GKE pods, serverless services, internet endpoints or on-premises endpoints.
  • Health check: probes each backend and removes failing ones from rotation.

Resource chain of a GCP proxy load balancer (Application Load Balancer shown)Forwarding ruleIP, port, schemeTarget proxyHTTPS + certificateURL maphost / path rulesBackend service ABackend service BMIG us-central1MIG europe-west1NEG (GKE pods)Health check35.191.0.0/16 probesPassthrough Network Load Balancers skip the proxy and URL map: forwarding rule → backend service → backendsand the backend sees the original client IP and terminates the connection itself
A proxy load balancer is a chain: forwarding rule, target proxy, URL map, backend services, backends, with health checks attached to each backend service. Passthrough load balancers have no proxy or URL map.

The product family

Google's current documentation groups the products into three families. The scheme names below are what you pass to --load-balancing-scheme.

ProductDeployment modesTrafficScheme
Application Load Balancer (external)Global; regional; classicHTTP, HTTPSEXTERNAL_MANAGED (global and regional)
Application Load Balancer (internal)Regional; cross-regionHTTP, HTTPSINTERNAL_MANAGED
Proxy Network Load Balancer (external)Global; regional; classicTCP, optional SSL offload (global)EXTERNAL_MANAGED; classic uses EXTERNAL
Proxy Network Load Balancer (internal)Regional; cross-regionTCPINTERNAL_MANAGED
Passthrough Network Load Balancer (external)Regional; global (Preview)TCP, UDP, ESP, GRE, ICMP and moreEXTERNAL; EXTERNAL_PASSTHROUGH for the global Preview
Passthrough Network Load Balancer (internal)Regional onlyTCP, UDP, SCTP, ICMP and moreINTERNAL

The classic modes are the older implementations. New deployments should use the global or regional managed modes, which support advanced traffic management that classic does not. Global external modes require Premium network tier; regional external modes can also use Standard tier, which keeps traffic on the public internet until it reaches the region.

Advertisement

Three data paths under the product names

The families differ because they run on different infrastructure, and the infrastructure explains the behaviour.

Google Front Ends (GFEs) carry the global external Application and proxy Network Load Balancers and the classic modes. A client connects to an anycast IP, the connection terminates at a GFE close to the client, and the GFE opens a separate connection to a backend in the best region with capacity. The backend sees a Google address as the source and gets the client address in X-Forwarded-For.

Managed Envoy proxies carry the regional external and internal Application Load Balancers, the regional proxy Network Load Balancers and the cross-region internal ones. Those proxies take their addresses from a proxy-only subnet that you create in each region and VPC; backend traffic arrives from that range, and firewall rules must allow it.

Passthrough load balancers have no proxy. The external passthrough Network Load Balancer is built on Google's Maglev, and the internal one is implemented in the Andromeda network virtualization stack. Packets are delivered to a backend VM with the original client source address, the VM terminates the connection, and responses go straight back to the client rather than through the load balancer. That is why only passthrough load balancers carry UDP and other non-TCP protocols, and why they cannot route on HTTP content or terminate TLS for you.

How to choose

  1. Does traffic need HTTP features such as path routing, header-based routing, TLS termination with managed certificates, Cloud CDN or Cloud Armor WAF rules? Use an Application Load Balancer.
  2. Is it non-HTTP TCP where you still want the load balancer to terminate connections, perhaps with TLS offload? Use a proxy Network Load Balancer.
  3. Is it UDP or another non-TCP protocol, or must the backend see the client's IP at the packet level, or do you want the lowest added latency? Use a passthrough Network Load Balancer.
  4. Internet-facing or internal? Internal load balancers get an address inside your VPC and are reachable from the VPC, peered networks and connected on-premises networks.
  5. Global or regional? Global gives one anycast IP and cross-region failover; regional keeps termination, TLS keys and traffic inside one region, which some data-residency rules require, and allows Standard tier for external traffic.

A common production layout uses a global external Application Load Balancer at the edge, with Cloud Armor and Cloud CDN attached, and regional internal Application Load Balancers or internal passthrough load balancers between tiers.

Worked example: a global HTTPS front end for two regions

Goal: serve www.example.com from managed instance groups in us-central1 and europe-west1, with health checks, a Google-managed certificate and one global IP. The commands create the chain from the backend outward.

# 1. Health check, and a firewall rule for probes AND proxied traffic from the GFEs
#    (IPv4 ranges from the current firewall-rules docs for instance-group backends)
gcloud compute health-checks create http web-hc --port=8080 --request-path=/healthz
gcloud compute firewall-rules create allow-lb-gfe --network=prod-vpc \
    --allow=tcp:8080 --source-ranges=35.191.0.0/16,130.211.0.0/22 --target-tags=web

# 2. Global backend service with two regional managed instance groups
gcloud compute backend-services create web-bs --global \
    --load-balancing-scheme=EXTERNAL_MANAGED --protocol=HTTP \
    --port-name=http --health-checks=web-hc
gcloud compute backend-services add-backend web-bs --global \
    --instance-group=web-us --instance-group-region=us-central1 \
    --balancing-mode=RATE --max-rate-per-instance=200 --capacity-scaler=1.0
gcloud compute backend-services add-backend web-bs --global \
    --instance-group=web-eu --instance-group-region=europe-west1 \
    --balancing-mode=RATE --max-rate-per-instance=200 --capacity-scaler=1.0

# 3. URL map, certificate, proxy, address and forwarding rule
gcloud compute url-maps create web-map --default-service=web-bs
gcloud compute ssl-certificates create web-cert --domains=www.example.com --global
gcloud compute target-https-proxies create web-proxy --url-map=web-map \
    --ssl-certificates=web-cert
gcloud compute addresses create web-ip --global
gcloud compute forwarding-rules create web-fr --global \
    --load-balancing-scheme=EXTERNAL_MANAGED --address=web-ip \
    --target-https-proxy=web-proxy --ports=443

Three things commonly surprise people. The managed certificate stays in a provisioning state until DNS for the domain points at the load balancer's IP, so create the address early and update DNS before expecting HTTPS to work. The firewall rule is mandatory: health-check probes come from 35.191.0.0/16 and proxied requests from the GFE ranges, and without it every backend is unhealthy and clients get errors. And the instance groups must define a named port http mapping to 8080, which is what --port-name refers to.

For GKE, you rarely type these commands: the Gateway or Ingress controller creates the same chain and uses NEG backends so the load balancer sends traffic directly to pods, as described in GKE architecture. The resource model is identical, which is what makes debugging a controller-created load balancer possible.

Regional and internal load balancers: the proxy-only subnet

Every Envoy-based load balancer needs a proxy-only subnet in its region and VPC before you create it. It is reserved for the managed proxies; you cannot put VMs in it. The documentation's minimum is a /26 (64 addresses) and its recommended starting size is a /23, because the number of proxies grows with traffic. Regional load balancers use the REGIONAL_MANAGED_PROXY purpose; cross-region internal ones use GLOBAL_MANAGED_PROXY.

# Required once per region and VPC before creating a regional Envoy-based load balancer
gcloud compute networks subnets create proxy-only-us \
    --purpose=REGIONAL_MANAGED_PROXY --role=ACTIVE \
    --region=us-central1 --network=prod-vpc --range=10.129.0.0/23
# Cross-region internal load balancers use --purpose=GLOBAL_MANAGED_PROXY instead

Plan these ranges as part of your IP address plan, alongside the rest of the VPC design, and add a firewall rule that allows the proxy-only range to reach backend ports. Forgetting that rule produces the same symptom as a missing health-check rule: the load balancer is created without error and never serves traffic.

How traffic is distributed

Each backend in a backend service has a balancing mode that defines its capacity: RATE (requests per second per instance or endpoint), UTILIZATION (instance-group CPU utilization) or CONNECTION (concurrent connections, for TCP). A capacity scaler between 0.0 and 1.0 multiplies that capacity; setting it to 0 drains a backend without removing it, which is the safe way to take a region out for maintenance.

The global external Application Load Balancer sends each request to the closest region that has healthy capacity and overflows to the next closest when a region is at its configured capacity. That makes the capacity numbers a traffic-engineering control, not just a limit: if max-rate-per-instance is set far above what an instance can actually serve, the load balancer keeps sending to an overloaded region instead of spilling over.

Health checks default to a 5-second interval and timeout, with two consecutive results needed to mark a backend healthy or unhealthy, so a failed backend leaves rotation within roughly 10 to 15 seconds. Point health checks at an endpoint that checks the instance's own readiness, not its dependencies, or one failing database takes every backend out at once.

Failure modes and operations

SymptomLikely causeCheck
All backends unhealthy right after creationFirewall does not allow the probe, GFE or proxy-only ranges; wrong port or pathBackend health in console; firewall rules on backend network tags
Intermittent 502s under loadBackend closes idle keepalive connections before the proxy doesSet server keepalive longer than the load balancer's documented value; many servers default to seconds
Requests cut off at a fixed durationBackend service timeout shorter than the requestRaise --timeout on the backend service for long requests or streaming
HTTPS not working on a new domainManaged certificate still provisioningDNS points at the LB IP; certificate status
Backend sees only Google or proxy IPsProxy load balancer, as designedRead X-Forwarded-For, or use a passthrough LB
One region overloaded while another idlesRATE capacity set too high to trigger overflowSet max rate from load tests; watch per-backend utilization

Enable logging on backend services, with a sample rate you can afford, and build dashboards on request count, latency and response code by backend. For Application Load Balancers, the log's status details field distinguishes responses the backend produced from errors the load balancer generated itself, such as failure to connect to a backend, which is the fastest way to tell an application bug from a network or health problem.

Trade-offs

Proxies give you L7 features, TLS termination and edge security, at the cost of hiding the client address in a header and adding a connection hop. Passthrough gives protocol freedom, client IPs and minimal overhead, but no content routing, and your VMs handle TLS and connection limits themselves. Global gives one IP and automatic cross-region failover; regional gives locality, data-residency control and Standard-tier pricing options. Envoy-based modes cost you a subnet per region to plan; GFE-based modes do not. Choose per traffic flow, not per company: most real systems run two or three types at once.

What to do next

  1. Inventory your traffic flows and classify each by protocol, internet or internal, global or regional, and whether backends need the client IP; map each to a product with the decision list above.
  2. Reserve proxy-only subnets in every region and VPC where you will run Envoy-based load balancers, and add them to your IP plan.
  3. Create firewall rules for health-check probes, GFE proxy ranges and proxy-only ranges before creating load balancers.
  4. Build the worked example in a test project, then drain one region with capacity-scaler 0 and confirm traffic fails over.
  5. Set RATE or UTILIZATION capacities from load tests, not guesses, so regional overflow actually triggers.
  6. Align server keepalive and backend service timeouts with the load balancer's documented values.
  7. Turn on backend service logging and alert on 5xx rate split by backend-generated versus load-balancer-generated errors.
Key takeaway: Every GCP load balancer is a chain of forwarding rule, optional target proxy and URL map, backend service, backends and health check, and the load-balancing scheme on the forwarding rule decides which product you have. Application Load Balancers do HTTP routing and edge security, proxy Network Load Balancers terminate TCP, and passthrough Network Load Balancers forward packets with client IPs intact. Choose global or regional deliberately, plan proxy-only subnets for Envoy-based modes, open firewalls for probes and proxies, set capacities that make overflow work, and read the logs to separate backend errors from load balancer errors.