Elastic Load Balancing is one AWS service with three current products that work at different layers. The Application Load Balancer understands HTTP and routes each request. The Network Load Balancer forwards TCP, UDP and TLS flows without reading them. The Gateway Load Balancer forwards raw IP packets through a fleet of security appliances and back. A fourth, the Classic Load Balancer, is legacy and not a choice for new systems.
Choosing wrongly is expensive to undo, because DNS names, static IPs, certificates and client behaviour all attach to the load balancer. This page explains how each one works inside, the defaults that cause incidents, how to drain targets during deployments, and a worked example that uses all three. General load-balancing algorithms and patterns across clouds are in Cloud Load Balancer.
The architecture
One family, three layers
| ALB | NLB | GWLB | |
|---|---|---|---|
| OSI layer | 7 (HTTP) | 4 (TCP, UDP, TLS) | 3 (IP packets) |
| Unit routed | Request | Flow | Flow of packets, both directions |
| Routing input | Host, path, headers, method, query, source IP | Flow hash of addresses and ports | Flow hash, through appliances |
| Targets | Instance, IP, Lambda | Instance, IP, ALB | Appliance instances or IPs |
| Static IPs | No; DNS name only | One per enabled AZ, Elastic IP optional | Reached via endpoints |
| Client source IP | In X-Forwarded-For header | Preserved by default for instance targets | Packets unchanged inside GENEVE |
| Cross-zone default | Always on at LB; can be off per target group | Off | Off |
All three share the same building blocks: a regional resource with nodes in each enabled Availability Zone, listeners that accept traffic on a port and protocol, rules or default actions that pick a target group, and target groups that hold targets and run health checks. Learn those four objects once and the differences become a matter of which layer each product reads.
Application Load Balancer: requests
An ALB terminates the client connection, parses HTTP and then opens or reuses its own connection to a target. Because it sees requests, it can route on host, path, headers, method, query strings and source IP, redirect, return fixed responses, authenticate users through an OIDC provider or Cognito, and split one rule across weighted target groups for canaries. It speaks HTTP/1.1, HTTP/2, gRPC and WebSockets to clients.
Each target group picks an algorithm: round robin (the default), least outstanding requests, or weighted random. Least outstanding requests suits uneven request cost. Weighted random supports automatic target weights, which reduce traffic to targets whose error rate is anomalous compared with their peers. Slow start ramps new targets linearly but works only with round robin. Stickiness uses a load-balancer cookie (AWSALB) or an application cookie, and requires cross-zone load balancing to be on for the target group.
Two behaviours surprise people. When no targets in a target group are healthy, the ALB fails open and sends traffic to all registered targets, on the theory that a broken health check is likelier than a dead fleet. And the ALB holds idle connections for its idle timeout (60 seconds by default): if a target's own keep-alive timeout is shorter, the target closes connections the ALB still considers usable and clients see intermittent 502 errors. Set the target's keep-alive longer than the ALB's idle timeout.
Network Load Balancer: flows
An NLB does not parse application data. Each node picks a target for a new flow using a hash of protocol, addresses and ports and then forwards every packet of that flow to the same target. Connection setup is cheap, latency is low, and the NLB exposes one IP per enabled Availability Zone, which you can pin with Elastic IPs. That makes it the answer for non-HTTP protocols, for clients that allow-list IPs, and for exposing a service through PrivateLink, which requires an NLB.
Flows have timeouts. TCP flows are tracked for 350 seconds of idleness by default, adjustable from 60 to 6,000 seconds; TLS listeners are fixed at 350 seconds and UDP flows at 120. After the timeout the next packet gets a reset. If you raise the TCP timeout above 350 seconds, raise the targets' connection tracking timeout to match, or the target side drops the flow first. Long-lived clients should send TCP keep-alives.
Client IP preservation is on for instance targets and for IP targets on UDP; IP targets on TCP and TLS default to off, in which case targets see the NLB's private address and can read the real client from proxy protocol v2 if you enable it. With preservation off, each unique target IP and port can take roughly 55,000 simultaneous connections before port allocation errors. With preservation on, a target that connects back to its own NLB can be routed to itself and fail, the hairpinning trap. Finally, security groups on an NLB must be attached at creation; an NLB created without them cannot gain them later.
Gateway Load Balancer: packets through appliances
A GWLB is a bump in the wire. You put security appliances (firewalls, intrusion detection, packet inspection) in a provider VPC behind a GWLB, publish it as an endpoint service, and create GWLB endpoints in the VPCs you want inspected. Route tables then send traffic to the endpoint as a next hop. The GWLB wraps each packet in GENEVE on UDP port 6081 and sends it to an appliance, which inspects it and returns it in the same encapsulation; the packet then continues to its original destination.
Inspection only works if both directions of a flow reach the same appliance, so the GWLB keeps flow stickiness on a 5-tuple by default, with 3-tuple and 2-tuple options for protocols that change ports. Appliances must speak GENEVE, answer health checks and scale as a fleet. Plan routing carefully: the endpoint must sit in a different subnet from the workloads it protects, and asymmetric routes that bypass the endpoint on return silently break stateful inspection.
Cross-zone, health and failover
Cross-zone load balancing decides whether a node in one zone may send traffic to targets in another. It is always on at the ALB level and can be turned off per target group; for NLB and GWLB it is off by default, and turning it on for an NLB adds inter-zone data transfer charges. With cross-zone off, each zone must hold enough healthy capacity for the traffic its node receives, because clients resolving the DNS name spread roughly evenly across zone IPs regardless of how many targets each zone has.
Health is layered. Targets pass or fail target-group health checks. Target-group health settings let you declare a minimum healthy count or percentage below which the load balancer fails over at DNS or routing level. NLB nodes also have zonal health: when a zone has no healthy targets, its IP is removed from DNS. Tune health checks to fail fast enough to matter and to test something real, such as a dependency-free readiness endpoint.
Draining and deployments
Deployments break on draining more often than on routing. When a target is deregistered it enters the draining state: no new requests or connections, while existing ones may finish for the deregistration delay, 300 seconds by default. Then it becomes unused and can be terminated. Three things must line up: the delay must exceed your longest normal request, the application must keep serving in-flight work after receiving SIGTERM, and the orchestrator must wait for draining before killing the process. ECS and Auto Scaling groups integrate with this lifecycle; custom scripts often do not.
For NLBs, a deregistered target that stays healthy can keep receiving traffic on existing connections; enable connection termination on deregistration if long-lived connections must move. NLB target groups also terminate established connections when a target turns unhealthy, which you can disable to let them close gracefully.
# Target group for ECS tasks, drained in 30 s instead of the 300 s default
aws elbv2 create-target-group --name api-tg --protocol HTTP --port 8080 \
--target-type ip --vpc-id vpc-0abc --health-check-path /ready
aws elbv2 modify-target-group-attributes --target-group-arn "$API_TG" --attributes \
Key=deregistration_delay.timeout_seconds,Value=30 \
Key=load_balancing.algorithm.type,Value=least_outstanding_requests
# ALB rule: /api/* goes to the api target group
aws elbv2 create-rule --listener-arn "$HTTPS_LISTENER" --priority 10 \
--conditions Field=path-pattern,Values='/api/*' \
--actions Type=forward,TargetGroupArn="$API_TG"
# NLB with security groups attached at creation (cannot be added later)
aws elbv2 create-load-balancer --name mqtt-nlb --type network --scheme internet-facing \
--subnets subnet-a subnet-b subnet-c --security-groups sg-0nlb
Worked example: an IoT platform
An IoT platform has a web console, a public REST and gRPC API, millions of devices speaking MQTT over TLS, and a compliance requirement that all traffic leaving the workload VPCs is inspected by a third-party firewall.
- Console and API: one ALB. Host and path rules send
console.example.comto an instance target group and/api/*to ECS tasks registered as IPs, with least outstanding requests because API calls vary in cost. A weighted forward action carries a 5 percent canary. - Devices: one NLB. MQTT is not HTTP, and device firmware allow-lists IPs, so the NLB gets three Elastic IPs and a TLS listener. The idle timeout is fixed at 350 seconds on TLS listeners, so devices must send MQTT pings well inside it. Security groups are attached at creation.
- Egress inspection: GWLB. Workload subnets route 0.0.0.0/0 to a GWLB endpoint; the firewall fleet sits behind the GWLB in a security VPC, with 5-tuple stickiness.
- Static entry for the API. If partners need fixed IPs for the ALB, either place it as a target behind an NLB or front it with Global Accelerator.
Failure modes
- Intermittent 502s. Target keep-alive shorter than the ALB idle timeout.
- Silent resets on long connections. NLB idle timeout reached with no keep-alives, or target conntrack shorter than a raised NLB timeout.
- Uneven zones. NLB cross-zone off with fewer targets in one zone; that zone's targets overload while others idle.
- Fail-open masking. A broken ALB health check marks everything unhealthy, the ALB fails open, and real failures pass unnoticed. Alarm on healthy host count, not just errors.
- Errors on every deploy. Processes killed before draining ends, or a delay shorter than real requests.
- No security groups. An NLB created without them cannot get them; recreate it.
- Broken inspection. Asymmetric routes around a GWLB endpoint drop stateful flows.
Cost and trade-offs
Cost follows the layer. ALBs are billed per hour plus load balancer capacity units, measured on the highest of new connections, active connections, processed bytes and rule evaluations, so many rules and small requests cost more than few rules and large ones. NLBs use their own capacity units on connections or flows and bytes, and cross-zone traffic adds data transfer. GWLB adds per-endpoint charges and the appliances themselves. Check the pricing page for your region rather than reusing numbers from old designs.
Functionally, the ALB buys request-level control and costs a little latency and no static IP. The NLB buys static IPs, any TCP or UDP protocol and very high connection rates, and gives up content-aware routing. The GWLB solves exactly one problem, transparent inspection, and should not be used for anything else.
What to do next
- Inventory your load balancers and replace any Classic Load Balancer with an ALB or NLB.
- For every ALB target, confirm keep-alive is longer than the ALB idle timeout.
- Set deregistration delay per target group to slightly above your longest normal request and test a deploy under load.
- For NLBs, record whether security groups exist and whether cross-zone is on, and size each zone for its share of traffic if it is off.
- Alarm on healthy host count per target group and per zone.
- Where clients hold long connections, set keep-alives inside the idle timeout.
- If you inspect traffic, verify symmetric routing through each GWLB endpoint.