Network address translation rewrites the addresses and ports in packet headers as they cross a boundary. It exists because there were never enough IPv4 addresses for every device, and it became so universal that almost every connection on the internet passes through at least one translator: the home router, the cloud NAT gateway in front of a private subnet, the carrier's large-scale NAT, the Kubernetes node rewriting a Service address.

Most of the time NAT is invisible. When it is not, the symptoms are confusing: connections that fail only under load, idle database sessions that hang instead of erroring, peers that cannot reach each other, and logs full of the gateway's address instead of the client's. This article builds NAT from the packet up, so each of those symptoms has an obvious cause, and then gives the arithmetic and configuration you need to run it.

Advertisement

What a translator rewrites

A transport connection is identified by a five-tuple: protocol, source address, source port, destination address and destination port. A NAT changes some of those fields on the way out and reverses the change on the way back. There are three basic forms.

  • Source NAT (SNAT) rewrites the source address of outbound packets, usually from a private range such as 10.0.0.0/8 to a public address. Masquerade is SNAT that uses whatever address the outgoing interface currently has, for links where it changes.
  • Destination NAT (DNAT) rewrites the destination of inbound packets, publishing an internal server on a public address and port. Port forwarding on a home router and most layer-4 load balancers are DNAT.
  • NAPT, network address and port translation, is SNAT that also rewrites the source port so that many private hosts can share one public address. This is what people almost always mean by NAT.

Rewriting a header has knock-on work. The IP header checksum and the TCP or UDP checksum, which covers a pseudo-header containing the addresses, must be recomputed. ICMP errors carry a copy of the original packet's header inside their payload, so a translator must rewrite that inner copy too, or path MTU discovery and traceroute break. Protocols that write addresses into their payload, such as FTP's active mode and SIP, need an application-level helper or they fail behind NAT.

A packet-level walkthrough

Source NAT with port translation (NAPT): one public address, many private flowsHost A10.0.1.15:51000Host B10.0.1.22:51000NAT gatewaypublic 203.0.113.7src rewrittenAPI server198.51.100.20:443Translation table (one entry per flow)inside srcoutside src (after NAT)destinationstate10.0.1.15:51000203.0.113.7:40001198.51.100.20:443 tcpESTABLISHED10.0.1.22:51000203.0.113.7:40002198.51.100.20:443 tcpESTABLISHEDBoth hosts used source port 51000; the NAT must pick different outside ports to keep replies apart.Replies to 203.0.113.7:40002 are matched against the table and rewritten back to 10.0.1.22:51000.No entry, no way back in: unsolicited inbound packets are dropped, which is why NAT feels like a firewall.
NAPT in one picture. Two private hosts happen to choose the same source port; the translator gives each flow a distinct outside port and keeps one table entry per flow to reverse the rewrite.

Follow host B's request. It sends a TCP SYN from 10.0.1.22:51000 to 198.51.100.20:443. The packet reaches the gateway, which finds no existing entry for that five-tuple, so this is a new flow. It chooses an outside port that is free for this destination, 40002, rewrites the source to 203.0.113.7:40002, records the pair of tuples in its table and forwards the packet.

The server sees a connection from 203.0.113.7:40002 and replies to it. The reply arrives at the gateway, which looks up the destination tuple, finds the entry, rewrites the destination back to 10.0.1.22:51000 and forwards it inside. Every later packet of the connection, in both directions, takes the same fast path through the table. When the connection closes, or its idle timer expires, the entry is removed and the outside port becomes reusable.

Two consequences follow directly. First, the table is state, so a NAT is a stateful device that can run out of entries and loses every connection if it restarts without replicating them. Second, an inbound packet that matches no entry has nowhere to go, so it is dropped. That is the property people mistake for security.

Advertisement

Mapping and filtering behaviour

How a NAT chooses and reuses outside ports matters enormously to anything peer-to-peer. RFC 4787 describes the behaviour in two independent dimensions.

DimensionEndpoint-independentAddress-dependentAddress and port-dependent
Mapping: when is the same outside port reused?Same inside tuple always maps to the same outside port, whatever the destinationReused only for the same destination addressReused only for the same destination address and port
Filtering: who may send packets in?Anyone, once the mapping existsOnly addresses the inside host has contactedOnly exact address and port pairs contacted

RFC 4787 recommends endpoint-independent mapping, because it lets a host learn its public address and port from one server, using STUN, and hand it to a peer who can then reach it. A NAT with address and port-dependent mapping, historically called symmetric NAT, gives every destination a different outside port, so the address learned from the STUN server is useless to the peer, and traffic must be relayed. The price of endpoint-independent mapping is capacity: an outside port stays tied to one inside flow for every destination, so one public address supports far fewer concurrent flows, which is why cloud egress gateways usually map per destination. The NAT traversal article covers STUN, TURN and ICE in detail.

Hairpinning is the related case of two hosts behind the same NAT reaching each other through its public address. A NAT that supports it rewrites both source and destination and sends the packet back inside; one that does not silently drops it, which is why an internal client sometimes cannot reach a service by its public name while external clients can.

NAT on Linux

Linux implements NAT in netfilter on top of connection tracking. Only the first packet of a flow traverses the NAT rules; the decision is stored in the flow's conntrack entry, and every later packet is rewritten from that entry. That is why a rule change does not affect existing connections, and why conntrack table exhaustion breaks NAT. The conntrack article covers the table itself. The nftables configuration for a small gateway looks like this:

table ip nat {
    chain prerouting {
        type nat hook prerouting priority dstnat; policy accept;
        # DNAT: publish an internal web server on the gateway's public address
        iifname "eth0" tcp dport 443 dnat to 10.0.1.30:8443
    }
    chain postrouting {
        type nat hook postrouting priority srcnat; policy accept;
        # SNAT: rewrite private sources to the fixed public address
        oifname "eth0" ip saddr 10.0.0.0/16 snat to 203.0.113.7
        # or, when the public address is dynamic (DHCP, PPPoE):
        # oifname "eth0" ip saddr 10.0.0.0/16 masquerade
    }
}

SNAT belongs in the postrouting hook, after the routing decision has chosen the outgoing interface, and DNAT in prerouting, before routing, so the rewritten destination is the one routed. Forwarding must be enabled and the filter rules must allow the forwarded traffic; NAT rules only rewrite. To see what the kernel is doing:

# Is forwarding on? NAT does nothing for routed traffic without it.
sysctl net.ipv4.ip_forward

# Inspect live translations: original tuple, then the reply tuple after NAT
conntrack -L -p tcp --dport 443 | head
# tcp 6 431999 ESTABLISHED src=10.0.1.22 dst=198.51.100.20 sport=51000 dport=443
#     src=198.51.100.20 dst=203.0.113.7 sport=443 dport=40002 [ASSURED]

# Table pressure and insert failures (drops show up here first)
conntrack -C ; sysctl net.netfilter.nf_conntrack_max
conntrack -S | grep -E 'insert_failed|drop'

Port exhaustion: the arithmetic

On translators that map per destination, as the AWS figures below imply, an outside port can be reused for different destinations, because the full five-tuple still differs. The binding limit is then per public address per destination: roughly the 64,000 or so usable ports, minus whatever the implementation reserves. AWS documents its NAT gateway as supporting up to 55,000 simultaneous connections from each of its IPv4 addresses to each unique destination, where a destination is the combination of IP address, port and protocol. A gateway can hold up to eight addresses, and when the limit is hit new connections fail and the ErrorPortAllocation CloudWatch metric rises.

Worked example: a fleet of 300 pods calls one third-party API at a single address on port 443. Each pod keeps a connection pool of 100 and churns connections because the client library closes them after each batch. Live connections are 300 x 100 = 30,000, comfortably inside one address's 55,000. But a closed connection does not necessarily free its mapping at once: the client holds TIME_WAIT for 60 seconds on Linux, and a translator typically keeps a closing entry for a while before reusing its port. If each pod opens 10 new connections per second and closing entries linger for about as long as TIME_WAIT, 300 x 10 x 60 = 180,000 mappings exist at once, more than three addresses' worth. The fix is not a bigger gateway first; it is connection reuse. Keep-alive pools that hold connections open cut the churn to almost nothing, and then add secondary addresses for real headroom.

Google Cloud NAT makes the per-VM side explicit: each VM gets a block of ports, with a default minimum of 64 per VM when dynamic port allocation is off and 32 when it is on. A busy VM talking to one destination exhausts 64 ports quickly, so either enable dynamic allocation with a sensible maximum or raise the minimum for those workloads.

Carrier-grade NAT and Kubernetes

Many mobile and broadband providers place subscribers behind a second, large-scale NAT, often using the shared address space 100.64.0.0/10 that RFC 6598 reserved for exactly this. Your home router translates once, the carrier translates again, and each subscriber receives a block of ports on a shared public address. For a server operator the consequences are that thousands of unrelated users can share one source address, so rate limiting or banning by IP punishes innocent users, and that identifying a user from a log requires the address, port and precise timestamp.

Kubernetes uses NAT in both directions. A ClusterIP Service is a virtual address that kube-proxy or an eBPF data plane DNATs to a chosen pod. Pod egress to destinations outside the cluster is usually masqueraded to the node's address. For traffic arriving through a load balancer, setting externalTrafficPolicy: Local avoids an extra SNAT hop between nodes and preserves the client's address, at the cost of only sending traffic to nodes that run a ready pod.

Timeouts and long-lived connections

A NAT cannot know whether an idle flow is dead, so it guesses with timers. RFC 4787 says UDP mappings should last at least two minutes and recommends five; RFC 5382 says an established TCP mapping should not expire in less than 2 hours and 4 minutes. Real devices are often shorter. AWS NAT gateways time out idle connections after 350 seconds and answer any later packet on them with a reset. Linux's default established-TCP conntrack timeout is five days.

When a timer expires on a connection both ends still believe is open, the next packet finds no mapping. Depending on the device it is dropped silently or answered with a reset. A database connection pool idle over lunch then hangs on the first query until the client's own timeout fires. The fix is to keep the flow active more often than the shortest timer on the path: set TCP keepalive on the client below it, for example 300 seconds idle before probes behind an AWS NAT gateway, or configure the pool to validate or recycle idle connections. For UDP protocols, send an application keepalive every 20 to 30 seconds. The TCP architecture article explains the keepalive knobs.

NAT is not a firewall, and IPv6 needs neither

The drop-unsolicited-inbound behaviour comes from the stateful table, not from translation. A stateful firewall with a default-deny inbound policy gives exactly the same protection without rewriting anything, and it is what you should rely on. Conversely, DNAT rules, UPnP and hairpinning open paths that a NAT-as-security mindset forgets about.

IPv6 provides enough addresses that translation for address sharing is unnecessary; hosts get globally unique addresses and a stateful firewall provides the protection. NAT survives in IPv6 networks mainly as NAT64, which lets IPv6-only clients reach IPv4-only servers, paired with DNS64 to synthesise addresses. The IPv6 architecture article covers that design.

Failure modes

SymptomCauseFix
New connections fail under load to one destinationPort exhaustion per destinationConnection reuse; more public addresses
Idle DB connections hang, then time outNAT idle timer shorter than pool idle timeTCP keepalive below the timer; pool validation
Internal clients cannot reach the public nameNo hairpinningSplit-horizon DNS or enable hairpin NAT
Peer-to-peer calls always relayedAddress and port-dependent mappingEndpoint-independent mapping; TURN as fallback
Random drops on a busy Linux gatewayconntrack table fullRaise nf_conntrack_max; shorten timeouts
Rate limiter blocks many users at onceUsers share a CGNAT addressLimit per account or token, not per IP

What to do next

  1. Draw every translator between your services and their dependencies: home or office router, cloud NAT gateway, Kubernetes Services, carrier NAT.
  2. For each cloud NAT gateway, chart the port-allocation error metric and active connection count, and alert before limits.
  3. Find the shortest idle timeout on each path and set client TCP keepalive or pool validation below it.
  4. Check outbound clients reuse connections; count new connections per second per destination.
  5. On Linux gateways, watch conntrack -S for insert failures and size nf_conntrack_max to peak flows.
  6. Replace any reliance on NAT for protection with an explicit default-deny stateful firewall policy.
  7. Where you can, enable IPv6 for egress to remove translation from the path entirely.
Key takeaway: NAT rewrites addresses and ports and remembers each rewrite in a table, and every NAT problem is a property of that table: entries are limited per public address and destination, they expire on timers that may be shorter than your idle connections, their mapping behaviour decides whether peers can connect, and their absence drops unsolicited traffic. Reuse connections, keep flows alive below the shortest timer, size and monitor port allocation, use a real firewall for protection, and prefer IPv6 where translation can disappear.