Cloud NAT lets Google Cloud resources without external IP addresses open connections to the internet, or with Private NAT to other private networks, while nothing outside can open connections to them. It is the standard way to give private VMs, GKE nodes and serverless workloads outbound access for package mirrors, vendor APIs and model hubs.
Setting it up takes two commands. Running it well means understanding one resource: source ports. Every connection through NAT occupies a NAT IP and port tuple. How many tuples each VM gets, how long they stay held after a connection closes, and how many NAT IPs you own decide whether a busy service works or silently drops packets. This article explains the mechanics, the timeouts, sizing with real numbers, GKE specifics and how to diagnose drops. Defaults below were checked against Google's documentation on 2026-10-04.
The VPC model around NAT is covered in GCP VPC, and private access to Google services in Private Google Access and Private Service Connect.
What Cloud NAT is
Cloud NAT is a distributed, software-defined service. It is not a gateway VM or appliance, so there is no instance to scale, patch or lose. A gateway belongs to one region and one VPC network. It is configured on a Cloud Router, which acts only as its control plane; packets do not flow through the router. A gateway serves the subnet ranges you select in that region, and supported sources include Compute Engine VMs, GKE nodes and pods, Cloud Run through Direct VPC egress or Serverless VPC Access, and App Engine standard.
There are two types. Public NAT, the default, translates private source addresses to a set of shared external IPs for internet access. Private NAT translates private addresses for traffic to connected private networks: other VPCs, on-premises sites or other clouds. Both allow response packets for connections a source opened and drop unsolicited inbound connections. One routing detail surprises people: traffic to Google APIs and services goes through Private Google Access even when the VM uses Public NAT, so Cloud Storage traffic does not consume NAT ports.
Ports, tuples and destinations
Each NAT IP offers 64,512 TCP source ports and 64,512 UDP source ports. The gateway hands each VM a block of them. A connection through NAT is identified by a five-tuple: NAT IP, NAT source port, and the destination IP, destination port and protocol. Because the destination is part of the tuple, a VM can reuse the same NAT port for connections to different destinations. The limit that bites is therefore per destination: a VM with 1,024 ports can hold 1,024 simultaneous connections to each distinct destination IP, port and protocol combination.
This is why NAT problems show up for one vendor API and not others. A service that talks to api.vendor.example:443 behind one or two IPs puts all its connections on the same few destination 3-tuples. The metric that measures this is compute.googleapis.com/nat/port_usage: the maximum number of connections from a VM to a single destination IP and port. It is the number to size against, not the VM's total connection count.
Static and dynamic port allocation
Ports are allocated to VMs in one of two ways.
| Static allocation | Dynamic port allocation (DPA) | |
|---|---|---|
| Default for | Public NAT | Private NAT |
| Minimum ports per VM | 64 by default | 32 by default; a power of 2 from 32 to 32,768 |
| Maximum ports per VM | Same as minimum | A power of 2 above the minimum, at most 65,536 (the default) |
| Behaviour | Fixed block per VM | Starts at the minimum and doubles as the VM nears exhaustion |
| Endpoint-independent mapping | Optional, off by default | Not allowed |
Static allocation is simple and predictable, but every VM holds its full block even when idle, and a busy VM cannot borrow. DPA gives quiet VMs few ports and busy ones many, which suits mixed fleets and GKE nodes. Growing takes time, though, so a VM that bursts from idle to thousands of connections can drop packets while its allocation doubles.
Endpoint-independent mapping (EIM) makes a VM's internal address and port map to the same NAT address and port whatever the destination, which some peer-to-peer and UDP protocols need. It reduces port reuse, can cause drops of its own, and rules out DPA. Leave it off unless a protocol needs it.
NAT IPs are either auto-allocated, where Google adds and removes addresses as VMs and port reservations change, or manual, where you name reserved static addresses. Use manual allocation whenever a partner allowlists your egress IPs. Manual IPs can be drained: new connections stop using them while existing ones finish. Switching between automatic and manual allocation breaks active connections.
Timeouts and the connection-rate ceiling
A tuple is not free the moment a connection ends. The gateway keeps mappings for set times, and these decide how fast ports can be reused.
| Timeout | Default | What it controls |
|---|---|---|
| UDP mapping idle | 30 s | How long an idle UDP mapping is kept |
| TCP established idle | 1,200 s (20 min) | How long an idle open connection keeps its mapping |
| TCP transitory idle | 30 s | Half-open connections |
| TCP TIME_WAIT | 30 s for gateways created after 2026-09-30; 120 s for older ones | How long a fully closed connection holds its tuple |
| ICMP mapping idle | 30 s (Public NAT) | Idle ICMP mappings |
Check what an existing gateway uses: gcloud compute routers nats describe NAT --router=ROUTER --region=REGION reports effectiveTcpTimeWaitTimeoutSec. Google recommends keeping TIME_WAIT at 15 seconds or more, and notes that a tuple may need up to 30 seconds beyond the configured transitory or TIME_WAIT value before it is reused.
This gives a ceiling on how fast a VM can open new connections to one destination: ports divided by (connection lifetime + TIME_WAIT + up to 30 s). With the static default of 64 ports, one-second connections and a 120 s TIME_WAIT, that is about 0.4 new connections per second to one destination, a limit a busy service hits instantly. The established idle timeout matters the other way: a pooled connection idle for more than 20 minutes loses its mapping, and the next packet fails unless the client sends keepalives.
Configuring a gateway
A Public NAT gateway with manual IPs, DPA and logging of errors:
REGION=us-central1
gcloud compute addresses create nat-ip-1 nat-ip-2 --region=$REGION
gcloud compute routers create nat-router --network=prod-vpc --region=$REGION
gcloud compute routers nats create prod-nat \
--router=nat-router --region=$REGION \
--nat-external-ip-pool=nat-ip-1,nat-ip-2 \
--nat-custom-subnet-ip-ranges=gke-subnet:ALL,batch-subnet \
--enable-dynamic-port-allocation \
--min-ports-per-vm=256 --max-ports-per-vm=4096 \
--tcp-time-wait-timeout=30s \
--enable-logging --log-filter=ERRORS_ONLYThe subnet syntax trips people up. SUBNET:ALL covers the primary range and every secondary range; a bare SUBNET covers only the primary range; SUBNET:RANGE_NAME covers one secondary range and not the primary. --nat-all-subnet-ip-ranges covers everything in the region. If a source's address is not in a covered range, its traffic is not translated and it has no internet access, with no error at the source. The log filter can be ALL, ERRORS_ONLY or TRANSLATIONS_ONLY; logging is off by default.
Worked example: sizing ports and IPs
Size from the busiest destination, not the average. For each VM, estimate the peak rate of new connections to that destination and how long each lasts, then compute ports and IPs:
import math
NAT_PORTS_PER_IP = 64_512
REUSE_SLACK_S = 30 # tuples may need up to 30 s beyond the timeout before reuse
def ports_per_vm(new_conn_per_s, lifetime_s, time_wait_s):
need = new_conn_per_s * (lifetime_s + time_wait_s + REUSE_SLACK_S)
return max(32, 2 ** math.ceil(math.log2(need))) # DPA values are powers of 2
def nat_ips(vms, ports):
return math.ceil(vms * ports / NAT_PORTS_PER_IP)
for tw in (30, 120):
ports = ports_per_vm(new_conn_per_s=50, lifetime_s=2, time_wait_s=tw)
print(tw, ports, nat_ips(120, ports))
# 30 4096 8
# 120 8192 16The worked case is 120 GKE nodes, each opening 50 short connections a second to one vendor API at peak. With a 30 s TIME_WAIT, each node needs about 3,100 tuples to that destination, rounded to 4,096 ports, and the fleet needs 8 NAT IPs if every node peaks at once. An older gateway still on 120 s needs 8,192 ports a node and 16 IPs. Now fix the client instead: with HTTP keep-alive and a pool of 200 long-lived connections per node, 256 ports a node is plenty and one IP covers all 120 nodes. Connection reuse is almost always cheaper than more NAT IPs, and it also avoids the TLS handshake on every call.
With static allocation the arithmetic runs the other way: at the default 64 ports, one IP serves 1,008 VMs, and each VM gets just 64 tuples per destination.
GKE and serverless egress
For GKE, the VM that receives ports is the node, so every pod on a node shares that node's allocation. A node running 60 pods that each call the same API needs 60 times the ports of one pod. That makes DPA, or a generous static minimum, the usual choice for clusters. Make sure the gateway covers the ranges your packets actually leave with, which for VPC-native clusters means the node primary range and the Pod secondary range; SUBNET:ALL or --nat-all-subnet-ip-ranges does that. Private clusters depend on Cloud NAT for any internet egress, including image pulls from registries outside Google. Cluster networking in general is covered in GKE architecture, and NAT for Shared VPC, where the gateway lives in the host project, in Shared VPC, in depth.
Monitoring and diagnosing drops
Four metrics, sampled every 60 seconds, tell you nearly everything; the first three are per VM:
nat/allocated_ports: ports currently assigned to the VM.nat/port_usage: the most connections the VM has to one destination IP and port. Alert when it approaches allocated ports.nat/dropped_sent_packets_countwith areasonlabel:OUT_OF_RESOURCESmeans the VM ran out of tuples;ENDPOINT_INDEPENDENCE_CONFLICTmeans an EIM conflict.nat/nat_allocation_failed, per gateway: the gateway could not give some VM its ports, usually because there are too few NAT IPs. The console reports how many more IPs you need.
Symptoms of exhaustion are intermittent: connect timeouts to one external service under load, while everything else works. Error-only logging records each dropped allocation with source, NAT IP and destination fields, which pinpoints the VM and destination. A dropped-packet alert with no matching application errors usually means retries are hiding it, and latency is paying.
Failure modes
- Port exhaustion on one destination. Short-lived connections to one API exceed per-VM tuples. Reuse connections, raise the minimum ports or enable DPA.
- Not enough NAT IPs.
nat_allocation_failedis true and new VMs get no ports. Add manual IPs or lower per-VM minimums. - Uncovered ranges. Pods or a new subnet are not in the gateway's ranges, so they have no egress and no error message.
- Idle pooled connections die. Connections idle past 20 minutes lose their mapping. Send keepalives or shorten pool idle times.
- Allowlist breakage. Auto-allocated IPs change; a partner's allowlist stops matching. Use manual, reserved addresses and drain before removing one.
- DPA burst lag. A node that jumps from idle to heavy load drops packets while its allocation doubles. Set the minimum near steady-state usage.
Trade-offs
Compared with external IPs on every VM, Cloud NAT keeps instances unreachable from the internet and gives a small, stable set of egress addresses, at the cost of shared port capacity to plan and monitor. Compared with a self-managed proxy fleet, it removes instances to run but offers no request-level inspection; when you need URL filtering, put a forward proxy such as Secure Web Gateway in front, which Cloud NAT supports as an endpoint type. Private Google Access, not NAT, is the right path for Google APIs. Behind a global external load balancer, inbound traffic never uses NAT at all; see GCP load balancing.
What to do next
- List every external destination your private workloads call, and find the busiest per VM.
- Read
port_usageandallocated_portsfor your top VMs, and alert on usage near allocation and on anydropped_sent_packets_count. - Check
effectiveTcpTimeWaitTimeoutSecon existing gateways; older ones may still hold tuples for 120 s. - Turn on connection pooling and keepalives in clients that call the same API repeatedly.
- Run the sizing script with your peak rates, then choose static or dynamic allocation and the number of NAT IPs.
- Switch to manual NAT IPs anywhere a partner allowlists you, and document the drain procedure.
- Confirm the gateway covers every primary and secondary range that needs egress, including GKE Pod ranges.