Instances in a private subnet usually still need outbound internet access, to pull packages, call third-party APIs or reach AWS endpoints that are not behind a VPC endpoint. On AWS, IPv4 egress from private subnets goes through network address translation, and you have two ways to provide it: a managed NAT gateway, or an EC2 instance you configure as a NAT yourself.
The choice looks simple: the gateway is managed and the instance is cheap. In practice it depends on traffic volume, availability targets, connection patterns and how much operational work you accept. This article explains how each option works, the limits that matter, the cost model with worked numbers, how to build a NAT instance correctly on a current OS, and how to decide. The surrounding VPC model, internet gateways and route tables, is covered in Amazon VPC.
What NAT does on the egress path
A private instance sends a packet to an internet address. Its subnet's route table sends 0.0.0.0/0 to the NAT. The NAT rewrites the source address to its own address and the source port to a free port, records the mapping, and forwards the packet. For a NAT gateway or NAT instance in a public subnet, that address is then translated once more by the internet gateway to the attached Elastic IP. Replies come back to the NAT's address and port, the NAT looks up the mapping, rewrites the destination back to the private instance, and delivers it. Connections initiated from the internet have no mapping and are dropped, which is the security property people rely on.
Two consequences follow. The NAT is stateful, so every active flow consumes a mapping and a source port, and mappings expire when flows go idle. And the NAT sits in one Availability Zone, so the route table that points at it decides which AZ failures your egress depends on.
The NAT gateway: a managed, zonal service
A NAT gateway is created in a subnet and is implemented with redundancy inside its Availability Zone. AWS documents its capacity: it supports 5 Gbps of bandwidth and scales automatically up to 100 Gbps, and processes one million packets per second, scaling automatically up to ten million, beyond which it drops packets. It supports TCP, UDP and ICMP, uses source ports 1024 to 65535, and can perform NAT64 for IPv6-only workloads when combined with DNS64 on the Route 53 Resolver.
Each IPv4 address on a NAT gateway supports up to 55,000 simultaneous connections to each unique destination, where a destination is the combination of destination IP, port and protocol. You can associate up to eight IPv4 addresses, one primary and seven secondary, raising the ceiling to 440,000 connections to a single destination; public NAT gateways are limited to two Elastic IPs by default until you request a quota increase. The port-pool mechanics behind that number are explained in managed NAT gateways.
Some things a NAT gateway cannot do. It cannot have a security group, so filtering happens on the instances and in the subnet's network ACL. You cannot route traffic into it through a VPC peering connection, or from a VPN or Direct Connect attached with a virtual private gateway; use a transit gateway for centralized egress, as described in AWS Transit Gateway. It does not do port forwarding and is not a bastion. A private NAT gateway variant, with no Elastic IP, translates between overlapping or private networks instead of to the internet.
Since November 2025 AWS has also offered a regional availability mode. A regional NAT gateway expands and contracts across the AZs where your workloads run, does not need a public subnet, and removes the per-AZ gateway and route table bookkeeping. It is priced per AZ it is active in, so it simplifies operations rather than reducing the hourly bill, and its CloudWatch metrics add an AvailabilityZone dimension. It is newer than the zonal design described in most guides; read the current documentation for IP address management modes before adopting it.
The NAT instance: an EC2 box you own
A NAT instance is an ordinary EC2 instance in a public subnet, with a public or Elastic IP, IP forwarding enabled, a masquerade rule and the EC2 source/destination check turned off so it can forward packets not addressed to itself. Private route tables point 0.0.0.0/0 at its instance ID or network interface.
AWS used to publish a NAT AMI, but it was built on Amazon Linux AMI 2018.03, which reached end of maintenance support on December 31, 2023. AWS now recommends migrating to NAT gateways, or building your own NAT AMI from a current Amazon Linux release if an instance suits you better. The build is short:
# NAT instance on current Amazon Linux (the old NAT AMI reached end of maintenance on 2023-12-31).
# Run as user data on an instance in a PUBLIC subnet with a public IP or Elastic IP.
dnf install -y iptables-services
systemctl enable --now iptables
cat > /etc/sysctl.d/90-nat.conf <<'EOF'
net.ipv4.ip_forward = 1
EOF
sysctl --system
IFACE=$(ip route show default | awk '{print $5}') # primary interface, e.g. ens5
iptables -t nat -A POSTROUTING -o "$IFACE" -j MASQUERADE
iptables -F FORWARD # drop the default REJECT rule
service iptables save
# From your workstation or automation, once:
# aws ec2 modify-instance-attribute --instance-id i-0abc... --no-source-dest-check
# aws ec2 replace-route --route-table-id rtb-0priv... \
# --destination-cidr-block 0.0.0.0/0 --network-interface-id eni-0nat...Its capacity is whatever the instance provides: network bandwidth and packets per second depend on the instance type, and burstable types run on credits, so a small instance that performs well in testing can throttle under sustained load. Its connection table is the kernel's conntrack table, bounded by nf_conntrack_max, and its timeouts are Linux defaults unless you tune them. In exchange you get things a NAT gateway lacks: security groups on the NAT itself, port forwarding, traffic inspection, logging of translations, and the option to double the box as a bastion, which is usually a bad idea.
Availability: zonal by default, either way
Both options fail per AZ. A NAT gateway is redundant within its AZ, but if that AZ fails, every route table pointing at it loses egress. The fix is architectural: one NAT gateway per AZ, with each private route table pointing at the gateway in its own AZ. Sharing one gateway across AZs saves an hourly charge but makes every AZ depend on one, and charges cross-AZ data transfer on every byte.
A single NAT instance is a single point of failure in every sense: instance failure, reboot, patching and AZ failure all stop egress. The common mitigation is an Auto Scaling group of size one per AZ whose instances claim the route on boot:
# Self-healing NAT instance: an Auto Scaling group of size 1 per AZ.
# On boot, the instance claims the route for its AZ's private route table.
TOKEN=$(curl -sX PUT http://169.254.169.254/latest/api/token \
-H "X-aws-ec2-metadata-token-ttl-seconds: 300")
IID=$(curl -s -H "X-aws-ec2-metadata-token: $TOKEN" \
http://169.254.169.254/latest/meta-data/instance-id)
aws ec2 modify-instance-attribute --instance-id "$IID" --no-source-dest-check
aws ec2 replace-route --route-table-id "$PRIVATE_RT" \
--destination-cidr-block 0.0.0.0/0 --instance-id "$IID"
# The instance role needs ec2:ModifyInstanceAttribute and ec2:ReplaceRoute,
# scoped by condition or resource to this route table. Outage = time to replace
# the instance and boot, typically minutes, during which egress from the AZ fails.This turns an indefinite outage into minutes, not zero, and existing connections are lost when the instance is replaced. Health checks should test forwarding, not only instance status. See Auto Scaling groups for health-check and replacement behaviour.
Idle timeouts and port exhaustion
A NAT gateway moves a connection to the idle state when it was not closed gracefully and has seen no activity for 350 seconds; the IdleTimeoutCount metric counts these. A client that then reuses the connection, such as a pooled database or HTTP connection, sends into a mapping that no longer exists and sees a reset or a hang. Set TCP keepalives or application pings below 350 seconds on long-lived pooled connections, or configure pools to discard idle connections sooner.
Port exhaustion shows up as ErrorPortAllocation greater than zero. It happens when a fleet opens many concurrent connections to the same destination, typically one API endpoint or one S3 address. Fixes, in order of preference: reuse connections with pooling and keepalive; send AWS service traffic through VPC endpoints; add secondary IPv4 addresses to the gateway; and split sources across several gateways. A NAT instance has the same arithmetic per public IP, plus its own conntrack table limit and Linux's long default established-connection timeout, so tune both if you run one at scale.
Cost, with numbers
A NAT gateway has two charges: an hourly charge per gateway and a per-gigabyte data processing charge on every byte it handles, in either direction and regardless of destination, on top of normal data transfer. At the time of writing, us-east-1 lists $0.045 per hour and $0.045 per GB, regional mode is $0.045 per hour per active AZ, and every public IPv4 address is $0.005 per hour. Check the pricing page for your Region.
| Scenario, 730 hours | Hourly | Processing | Total per month |
|---|---|---|---|
| 3 zonal gateways, 1 TB processed | 3 x 0.045 x 730 = $98.55 | 1,024 x 0.045 = $46.08 | about $145 plus $10.95 for 3 EIPs |
| 3 zonal gateways, 20 TB processed | $98.55 | 20,480 x 0.045 = $921.60 | about $1,020 plus EIPs |
| Same, 15 TB of it moved to an S3 gateway endpoint | $98.55 | 5,120 x 0.045 = $230.40 | about $329 plus EIPs |
The pattern is clear. At low volume the hourly charge dominates and a NAT instance on a small instance type can be cheaper. At high volume the processing charge dominates, and the biggest savings come from not sending traffic through any NAT: S3 and DynamoDB gateway endpoints have no charge, and interface endpoints, explained in AWS PrivateLink, have their own hourly and per-GB pricing that is often lower. A NAT instance avoids processing charges, but you pay for the instance, its EIP, your engineering time and the outages.
Worked example: choosing for three environments
A team runs three environments. Development moves a few gigabytes a day and tolerates an hour of egress loss: one NAT instance per VPC on a small current-generation instance, in an Auto Scaling group of one, is reasonable. Staging mirrors production topology to catch routing mistakes, so it uses NAT gateways, but only in the AZs it actually uses. Production moves 20 TB a month and has an availability target: one NAT gateway per AZ, an S3 gateway endpoint, which removes most of the processing charge in this example because the bulk of traffic is S3, and alarms on the metrics below.
# One NAT gateway per AZ, each private route table pointing at its own AZ's gateway (bash 4+).
declare -A PUBLIC_SUBNET=([a]=subnet-0pa [b]=subnet-0pb [c]=subnet-0pc)
declare -A PRIVATE_RT=([a]=rtb-0ra [b]=rtb-0rb [c]=rtb-0rc)
for AZ in a b c; do
EIP=$(aws ec2 allocate-address --domain vpc --query AllocationId --output text)
NAT=$(aws ec2 create-nat-gateway --subnet-id "${PUBLIC_SUBNET[$AZ]}" \
--allocation-id "$EIP" --query NatGateway.NatGatewayId --output text)
aws ec2 wait nat-gateway-available --nat-gateway-ids "$NAT"
aws ec2 create-route --route-table-id "${PRIVATE_RT[$AZ]}" \
--destination-cidr-block 0.0.0.0/0 --nat-gateway-id "$NAT"
done
# Keep S3 traffic off the NAT entirely: a gateway endpoint has no processing charge.
aws ec2 create-vpc-endpoint --vpc-id "$VPC" --vpc-endpoint-type Gateway \
--service-name com.amazonaws.us-east-1.s3 \
--route-table-ids "${PRIVATE_RT[@]}"
Operational checklist and failure modes
- Alarm on
ErrorPortAllocationabove zero,PacketsDropCountas a share of packets in,PeakPacketsPerSecondandPeakBytesPerSecondagainst the documented ceilings, and a sudden rise inIdleTimeoutCount. - Cross-AZ routing. Audit every private route table and confirm its NAT target is in the same AZ. This is the most common hidden availability and cost problem.
- MTU. NAT gateways support an 8,500-byte MTU, but internet paths are typically 1,500, so keep instance MTU at or below 1,500 for internet traffic.
- Return traffic over peering. Traffic from a NAT gateway into a peered VPC returns to the gateway automatically; use network ACLs if that path should not exist.
- NAT instance drift. Rebuild NAT instances from a pipeline-built image, patch them like any server, and test failover regularly.
What to do next
- Export every route table with
aws ec2 describe-route-tablesand confirm each private table targets a NAT in its own AZ. - Pull a month of
BytesOutToDestinationandBytesInFromDestinationper gateway and estimate your processing charge. - Use VPC flow logs to find the top destinations through each NAT; add gateway endpoints for S3 and DynamoDB and price interface endpoints for the rest.
- Create CloudWatch alarms for
ErrorPortAllocation,PacketsDropCountandIdleTimeoutCount. - Set keepalives below 350 seconds on pooled connections that traverse the NAT.
- If you still run the old NAT AMI, replace it with NAT gateways or a NAT instance built from a current Amazon Linux image.
- Evaluate regional mode in a non-production VPC if per-AZ gateway management is a burden.