Azure Load Balancer is not one product but a small family of layer 4 load balancers that share one data plane: the regional Standard load balancer in public and internal forms, a Global tier that spreads traffic across regions behind one anycast IP, and Gateway Load Balancer, which inserts firewalls and other network appliances into a flow without changing routes. The older Basic SKU was retired on 30 September 2025, so anything still described as "Basic" in an old template needs upgrading.
All of them work on flows, not requests. They do not terminate TLS, read HTTP headers or retry failed requests; they pick a backend for each new flow and forward packets with the client's address intact. That makes them fast and simple, and it moves a lot of responsibility onto you: health probes that mean something, idle timeouts that match your connection pools, and outbound port budgets that do not run out at peak. This article works through each member of the family, the mechanics that trip people up, and how to operate them.
The family at a glance
| Member | Scope | Frontend | Typical use |
|---|---|---|---|
| Standard public LB | One region, zone-redundant or zonal | Public IP | Internet-facing TCP/UDP services, outbound SNAT via outbound rules |
| Standard internal LB | One virtual network (reachable over peering) | Private IP | Tiers inside a VNet, HA ports in front of appliances, Private Link services |
| Global tier | Many regions | Static anycast public IP | One address for regional deployments, failover between regions |
| Gateway LB | Chained to a public frontend | Private IP | Transparent firewalls, IDS/IPS, packet inspection |
| Basic (retired) | Availability set | Public or private | Nothing new; migrate |
General concepts such as layer 4 versus layer 7 are in the cloud load balancer overview.
How a Standard load balancer forwards a flow
When the first packet of a new flow arrives at a frontend, the load balancer hashes the five-tuple (source IP, source port, destination IP, destination port, protocol) and picks a healthy backend from the pool. Every later packet of that flow hashes to the same backend. The TCP handshake is between the client and the VM, not with the load balancer, so the VM sees the real client IP and the load balancer keeps no connection state you can tune beyond the idle timeout.
So the unit of balance is a flow, not a request: one long-lived gRPC connection stays on one VM, and traffic from a few SNATed sources may spread unevenly. Changing pool membership remaps new flows, so returning clients may land elsewhere. Router-level ECMP is a good mental model. Standard is also closed by default: internet traffic reaches backends only if a network security group allows it.
Distribution modes and session persistence
The default five-tuple mode, shown as session persistence None and Default in the API, gives stickiness only within one connection. Two alternatives hash fewer fields: Client IP (SourceIP, a two-tuple of source and destination IP) and Client IP and protocol (SourceIPProtocol, a three-tuple). They keep all of a client's connections on one VM, which helps with a TCP control channel plus a UDP data channel, or with software that keeps session state in memory.
The cost is balance: everyone behind one corporate proxy looks like one client. Prefer stateless applications and the default mode.
Health probes that mean something
A load-balancing rule needs a health probe, and the probe decides which backends receive new flows. Probes come from 168.63.129.16 (IPv6 probes from a link-local address), and the AzureLoadBalancer service tag allows them by default in network security groups. Guest firewalls are another matter: a host firewall that drops that address marks every instance down.
Three protocols exist. A TCP probe succeeds if the handshake completes. HTTP and HTTPS probes send a GET to a path and count only a 200 as healthy; a 301, 404 or 500 is a failure. HTTP probes time out after 30 seconds. The default interval is 5 seconds in the portal and 15 seconds through the API and CLI, with a minimum of 5. The probe threshold, the number of consecutive results needed to change state, applies to TCP probes and to HTTP probes only when they time out; an explicit HTTP answer changes state immediately.
When one instance is marked down, its established TCP connections continue and only new flows go elsewhere. On Standard, even if every instance is down, established TCP flows continue (as long as the pool has more than one instance) while new flows fail. That is what makes a probe a drain switch: fail the probe on shutdown, wait for connections to finish, then stop the process.
# Health endpoint on :8081 that lets the instance drain itself before shutdown.
import signal, threading
from http.server import BaseHTTPRequestHandler, HTTPServer
draining = threading.Event()
signal.signal(signal.SIGTERM, lambda *_: draining.set())
def dependencies_ok() -> bool:
return True # e.g. local cache warm, DB pool has free connections
class Health(BaseHTTPRequestHandler):
def do_GET(self):
ok = self.path == "/healthz" and not draining.is_set() and dependencies_ok()
self.send_response(200 if ok else 503) # only 200 counts as up
self.end_headers()
HTTPServer(("0.0.0.0", 8081), Health).serve_forever()Probe local readiness, not shared dependencies, or a slow database takes every instance out at once.
Internal load balancers and HA ports
An internal load balancer has a private frontend for traffic inside the virtual network and across peering. Its special feature, the HA ports rule, balances all TCP and UDP flows on all ports; it exists for firewalls and other appliances, reached by pointing a user-defined route at the frontend.
With HA ports one probe represents the whole appliance, so probe the appliance itself, never a port it forwards to machines behind it, or one flapping backend can knock out every appliance.
Outbound: SNAT ports and outbound rules
A backend VM without its own public IP reaches the internet through SNAT. Each outbound connection needs a SNAT port, and each public IP offers 64,000. A TCP port can be reused for different destination IPs or ports, but not for two simultaneous connections to the same destination IP and port; UDP uses one port per destination IP. When an instance uses up its allocation, new connections to a busy destination fail even if the load balancer as a whole has spare ports.
Azure ranks outbound methods in order: a NAT Gateway on the subnet first (it takes precedence over everything else), then a public IP on the VM, then load balancer outbound rules, then implicit SNAT through a load-balancing rule, then default outbound access. New virtual networks created from 31 March 2026 use private subnets by default, so default outbound access is no longer something to rely on. Implicit SNAT allocates ports from a fixed table: 1,024 ports per instance for pools of 1 to 50, falling to 512, 256, 128, 64 and 32 for pools of up to 1,000, never above 1,024 per instance. It is not meant for production.
Outbound rules make the budget explicit: ports per instance = frontend IPs × 64,000 / instances. For scale sets, size for the maximum instance count rather than today's, or a scale-out will be blocked because there are no ports left to give new instances. For Azure PaaS services, use Private Link so the traffic never needs SNAT; for large or bursty egress, use a NAT gateway.
Worked example: sizing outbound for a scale set
A scale set runs 20 to 100 instances behind a public Standard load balancer and calls one payment API at a single IP on port 443. At peak each instance opens 25 new connections a second, and without connection reuse each lives about 30 seconds including teardown. That is about 750 ports held per instance for one destination.
Implicit SNAT gives 512 ports per instance at 51 to 100 instances: exhaustion at peak, seen as intermittent connect timeouts. One outbound IP sized for 100 instances gives 640: still short. Two give 1,280, which the CLI below allocates.
The better fix is in the client: a pool of 40 keep-alive connections per instance cuts the need to about 40 ports. Do both.
RG=web-prod; LB=web-lb
az network lb create -g $RG -n $LB --sku Standard \
--public-ip-address web-pip --frontend-ip-name fe-in --backend-pool-name be
# HTTP probe on a dedicated health port, every 5 seconds
az network lb probe create -g $RG --lb-name $LB -n health \
--protocol Http --port 8081 --path /healthz --interval 5
# Inbound rule: no implicit outbound SNAT, longer idle timeout, RST on idle
az network lb rule create -g $RG --lb-name $LB -n https \
--protocol Tcp --frontend-port 443 --backend-port 443 \
--frontend-ip-name fe-in --backend-pool-name be --probe-name health \
--idle-timeout 15 --enable-tcp-reset true --disable-outbound-snat true
# Explicit outbound: two frontend IPs = 128,000 ports = 1,280 each for 100 instances
for i in 1 2; do
az network public-ip create -g $RG -n web-out-pip$i --sku Standard
az network lb frontend-ip create -g $RG --lb-name $LB -n fe-out$i \
--public-ip-address web-out-pip$i
done
az network lb outbound-rule create -g $RG --lb-name $LB -n out \
--frontend-ip-configs fe-out1 fe-out2 --address-pool be --protocol All \
--outbound-ports 1280 --idle-timeout 15 --enable-tcp-reset true
Idle timeout and TCP reset
The load balancer forgets an idle flow after the idle timeout: 4 to 100 minutes on load-balancing and inbound NAT rules, 4 to 120 on outbound rules, 4 by default. By default the flow is dropped silently, so a pool holding idle connections for 10 minutes against a 4-minute timeout produces the classic "first request after a quiet period hangs" bug.
Enable TCP reset on idle on every rule so both ends get a reset and the pool discards the dead connection at once, then keep the pool's idle time below the timeout or use keepalives shorter than it.
The Global tier and Gateway Load Balancer
The Global tier is a public load balancer whose backend pool holds regional Standard public load balancers. Its static anycast IP is advertised from participating regions; traffic enters the nearest one and crosses Microsoft's backbone to the closest healthy regional deployment, checked every 5 seconds. Client IPs are preserved. Limits matter: frontends are public only, internal load balancers cannot be members, there are no outbound rules, and an existing regional load balancer cannot be upgraded to the Global tier. It is an alternative to DNS-based failover that avoids waiting for DNS caches.
Gateway Load Balancer puts appliances in the path transparently. You chain a Standard public frontend, or a VM's Standard public IP configuration, to a Gateway load balancer frontend; traffic to and from that endpoint is encapsulated in VXLAN and sent through the appliance pool first, with symmetric flow stickiness so both directions hit the same appliance. The backend pool uses up to two tunnel interfaces, external for traffic arriving from the internet and internal for traffic heading to the application. Its rules are HA ports only, it cannot be used as a next hop in a user-defined route, and it does not work with the Global tier. The appliances must support VXLAN.
When to use Application Gateway, Front Door or Traffic Manager instead
The family stops at layer 4. TLS termination, path routing, a WAF or per-request retries in a region is Application Gateway; the same at the global edge with caching is Front Door; DNS-level steering, including to endpoints outside Azure, is Traffic Manager. Choose the L4 family for non-HTTP protocols, UDP, lowest latency, client IP preservation or appliances in the path.
Observability and alerting
Azure Monitor carries the metrics. Data Path Availability is a synthetic probe through the platform; Health Probe Status is your probe's view. Data path healthy with probes failing means your application or NSG; both failing points at the platform. SNAT Connection Count by connection state shows failed outbound connections, and Used versus Allocated SNAT Ports shows headroom per instance.
- Alert when Data Path Availability for any frontend and port drops to zero, and when an instance's Health Probe Status averages below 100 for five minutes.
- Alert on any failed SNAT connections, split by backend IP.
- Alert at 75 percent and again at 90 percent of allocated SNAT ports per instance.
- Add a Resource Health alert for the Degraded and Unavailable states.
Failure modes
| Symptom | Likely cause | Fix |
|---|---|---|
| All instances down right after deployment | Guest firewall or NSG blocks 168.63.129.16, or the probe path returns a redirect | Allow the probe source; return 200 from the probe path |
| First request after a quiet spell hangs | Silent idle-timeout drop under a long-lived pool | TCP reset on idle; pool idle time below the timeout |
| Intermittent outbound connect timeouts at peak | SNAT exhaustion to one destination | Connection reuse, outbound rules sized for max instances, NAT Gateway, Private Link |
| One VM much hotter than the others | Few long-lived connections or source-IP persistence behind a NAT | Default mode, client-side connection spreading, more client connections |
| Whole service dark when the database is slow | Probe checks shared dependencies | Probe local readiness only |
What to do next
- Migrate any remaining Basic load balancers and public IPs to Standard.
- Give every rule a dedicated HTTP health endpoint that returns 200 only when the instance can serve, and wire it to a drain on shutdown.
- Turn on TCP reset on idle for every rule and align connection-pool idle times with the timeout.
- Set
--disable-outbound-snaton load-balancing rules and choose outbound explicitly: NAT Gateway first, outbound rules sized for maximum scale otherwise, Private Link for PaaS. - Create the data-path, probe-status, failed-SNAT and SNAT-usage alerts listed above.
- For multi-region services, compare the Global tier with DNS failover and test a region failure end to end.