AWS sells two services under the VPN name, and they solve different problems. Site-to-Site VPN joins a whole network, such as an office or a data centre, to your AWS network over IPsec tunnels across the internet. Client VPN lets individual people connect from laptops to resources inside a VPC, using an OpenVPN-based client and managed authentication. Most organisations end up running both, often meeting at the same Transit Gateway, and most of the problems they hit are in the seams: routes, address overlaps, MTU and authorization.

This article explains both from first principles, with the limits that shape designs taken from the AWS documentation, then works through a hybrid design for an organisation with offices and remote staff. It assumes you know the basics of VPC routing and builds on Transit Gateway. If you need dedicated bandwidth rather than internet tunnels, Direct Connect is the alternative, and Site-to-Site VPN is often its backup.

Advertisement

Site-to-Site VPN: the moving parts

A Site-to-Site VPN connection has three pieces. The customer gateway is an AWS record describing your on-premises device: its public IP and, for dynamic routing, its BGP autonomous system number. The AWS side terminates either on a virtual private gateway, attached to exactly one VPC, or on a Transit Gateway, which can route to many VPCs, other VPNs and Direct Connect. The VPN connection links the two and always consists of two IPsec tunnels, each with its own AWS public IP in a different Availability Zone.

Both tunnels should be configured on your device. AWS performs maintenance on one tunnel at a time, so a connection with only one tunnel configured goes down during routine work. Traffic from your network to AWS can use both tunnels; traffic from AWS back to you prefers one of them and fails over to the other. That asymmetry matters if a firewall on your side tracks sessions per tunnel: a flow that leaves through tunnel 1 and returns through tunnel 2 may be dropped, so either make the tunnels active-standby with BGP attributes or make the firewall tolerate asymmetric paths.

Two VPN services into one AWS networkRemote laptopsOpenVPN-based clientOffice / data centrecustomer gateway deviceClient VPN endpointauth, client CIDR, rulesS2S VPN connection2 IPsec tunnels, BGPAssociated subnetsone per AZ, ENIs, SGTransit Gatewayroute tables, ECMPWorkload VPCsprivate subnetsTLS, UDP 443tunnel 1tunnel 2SNAT to ENIVPC routeVPN attachmentattachmentsClient VPN puts users into a VPC; Site-to-Site VPN joins whole networks. Both meet at the Transit Gateway.Remote users reach the office through the same Transit Gateway if routes and authorization rules allow it.
Client VPN terminates users in associated subnets of a hub VPC; Site-to-Site VPN attaches whole networks to the Transit Gateway. Routes and authorization rules decide who reaches what.

Tunnels, bandwidth and ECMP

A standard tunnel carries up to 1.25 Gbps and up to 140,000 packets per second. Large Bandwidth Tunnels carry up to 5 Gbps and 400,000 packets per second, but only on connections attached to a Transit Gateway or Cloud WAN, not on a virtual private gateway; both tunnels of a connection must use the same size, the customer gateway needs a fixed IP, and accelerated VPN is not supported with them. These are ceilings, not guarantees: packet size, the TCP and UDP mix and internet conditions all reduce what you actually get, and small packets hit the packets-per-second limit before the bandwidth limit.

To go beyond one tunnel, use equal-cost multi-path routing on a Transit Gateway. ECMP spreads flows across tunnels that advertise the same prefixes, and it only works with dynamic, BGP-routed connections, never with static routes. AWS's own example is two connections with Large Bandwidth Tunnels and ECMP across all four tunnels for about 20 Gbps. ECMP balances per flow, so a single large transfer still rides one tunnel; parallel streams are how bulk copies benefit.

resource "aws_ec2_transit_gateway" "hub" {
  amazon_side_asn  = 64512
  vpn_ecmp_support = "enable"          # lets several tunnels share load when routing is dynamic
}

resource "aws_customer_gateway" "office" {
  bgp_asn    = 65010                   # the office router's ASN
  ip_address = "203.0.113.10"          # its fixed public IP
  type       = "ipsec.1"
}

resource "aws_vpn_connection" "office" {
  customer_gateway_id = aws_customer_gateway.office.id
  transit_gateway_id  = aws_ec2_transit_gateway.hub.id
  type                = "ipsec.1"
  static_routes_only  = false          # BGP: failover and ECMP need it
  tunnel1_inside_cidr = "169.254.10.0/30"
  tunnel2_inside_cidr = "169.254.10.4/30"
}

The Terraform above creates a hub Transit Gateway with ECMP enabled, a customer gateway for an office router, and a BGP-routed connection with explicit tunnel inside addresses. Inside CIDRs are /30 blocks from the link-local 169.254.0.0/16 range; choosing them yourself avoids collisions when the same router terminates tunnels to several connections.

Advertisement

Routing: static or BGP, and the limits

Static routing means you list the on-premises prefixes on the AWS side and AWS sends traffic for them to the connection. It is simple, and it is blind: if a tunnel's IPsec session stays up but the path behind it breaks, AWS keeps sending. BGP lets each side advertise prefixes and withdraw them on failure, gives faster and more precise failover, and is required for ECMP. Use BGP unless your device cannot do it.

Route limits are hard and they bite at scale. A customer gateway can advertise at most 100 dynamic routes to a connection on a virtual private gateway, and no more than 100 static routes can be configured there. On a Transit Gateway it can advertise 1,000. In the other direction, AWS advertises up to 1,000 routes to your device from a virtual private gateway and up to 5,000 from a Transit Gateway. None of these can be raised. If your data centre has hundreds of prefixes, summarise them before advertising, or terminate on a Transit Gateway.

When both Direct Connect and a Site-to-Site VPN advertise the same prefix to the same gateway, AWS prefers Direct Connect. That is what makes VPN a natural backup circuit. Test the failover by withdrawing the prefix from Direct Connect, not by assuming it will work.

MTU and MSS

IPsec adds headers, so the largest packet that fits inside a tunnel is smaller than a normal Ethernet frame. AWS supports an MTU of 1,446 bytes and an MSS of 1,406 bytes on Site-to-Site VPN, less with algorithms that use larger headers. Jumbo frames are not supported and neither is Path MTU Discovery. The symptom of getting this wrong is characteristic: ping and SSH work, but large HTTPS responses or database result sets hang. Clamp TCP MSS on the customer gateway to the documented value or lower, and set the tunnel interface MTU to match.

Client VPN: the moving parts

A Client VPN endpoint is a managed OpenVPN-compatible server. It has a client CIDR from which connected users get addresses, a server certificate in ACM, one or more authentication methods, and target network associations, which place network interfaces in subnets of one VPC. Users download a configuration file, connect with the AWS-provided client or another OpenVPN client, and their traffic enters the VPC through those interfaces.

Three authentication methods exist: mutual certificate authentication, Active Directory through AWS Directory Service, and SAML federation with an identity provider. SAML is usually the right choice for a workforce, because it reuses the company's MFA and joiner-and-leaver process; SAML federation covers the protocol. The self-service portal for downloading the configuration is not available to users who authenticate with mutual certificates.

Two separate controls decide what users can reach. Routes on the endpoint decide which destinations are sent through which associated subnet. Authorization rules decide which users, by group from the directory or the SAML assertion, may reach which destination CIDRs. Both must allow a flow. Forgetting authorization is the most common reason a new endpoint connects successfully but reaches nothing.

For IPv4, traffic leaving the endpoint is source-translated to the IP address of the association's interface. Your servers therefore see the endpoint, not the user's client address, so security groups on targets should reference the security group attached to the endpoint, and per-user attribution has to come from the endpoint's connection logs. IPv6 traffic is not translated.

Setting up a Client VPN endpoint

# 1. Endpoint: server cert in ACM, SAML IdP, logs, split tunnel
aws ec2 create-client-vpn-endpoint \
  --client-cidr-block 10.250.0.0/20 \
  --server-certificate-arn arn:aws:acm:eu-west-1:111122223333:certificate/SERVER-CERT-ID \
  --authentication-options Type=federated-authentication,FederatedAuthentication={SAMLProviderArn=arn:aws:iam::111122223333:saml-provider/Corp} \
  --connection-log-options Enabled=true,CloudwatchLogGroup=/vpn/client,CloudwatchLogStream=conn \
  --split-tunnel --transport-protocol udp --vpn-port 443 --session-timeout-hours 12 \
  --vpc-id vpc-0hub --security-group-ids sg-0vpnclients \
  --dns-servers 10.0.0.2

# 2. Associate one subnet per AZ (each association adds capacity and resilience)
aws ec2 associate-client-vpn-target-network --client-vpn-endpoint-id cvpn-endpoint-0abc --subnet-id subnet-0az1
aws ec2 associate-client-vpn-target-network --client-vpn-endpoint-id cvpn-endpoint-0abc --subnet-id subnet-0az2

# 3. Routes: which destinations go through the tunnel (per associated subnet)
aws ec2 create-client-vpn-route --client-vpn-endpoint-id cvpn-endpoint-0abc \
  --destination-cidr-block 10.20.0.0/16 --target-vpc-subnet-id subnet-0az1

# 4. Authorization: which groups may reach which destinations
aws ec2 authorize-client-vpn-ingress --client-vpn-endpoint-id cvpn-endpoint-0abc \
  --target-network-cidr 10.20.0.0/16 --access-group-id engineering

# 5. The file users import into the client
aws ec2 export-client-vpn-client-configuration --client-vpn-endpoint-id cvpn-endpoint-0abc --output text > corp.ovpn

The options have defaults worth knowing: UDP is the default transport and TCP the alternative, the port is 443 or 1194, split tunnel is off unless requested, and the maximum session length is 8, 10, 12 or 24 hours, with 24 the default. The client CIDR must be between /22 and /12, must not overlap the VPC or routes on the endpoint, and cannot be changed later, so plan it before creating the endpoint.

Several properties, including the client CIDR, authentication options and transport protocol, are fixed at creation. Others can be modified, but changes can take up to four hours to take effect, and changing the server certificate, DNS servers, split tunnel setting, routes, the revocation list, authorization rules or the port resets active connections. Batch changes and make them outside working hours.

Worked example: two offices and 1,500 remote staff

An organisation has a London office and a Pune data centre, 1,500 staff who work remotely at least part of the week, and workloads in six VPCs attached to a Transit Gateway in eu-west-1. The data centre advertises 40 prefixes; London advertises 12.

Site-to-Site. Each site gets a BGP connection to the Transit Gateway. Pune moves nightly database backups, so it uses Large Bandwidth Tunnels; London uses standard ones. Both are well inside the 1,000-route limit. MSS is clamped at 1,406 on both routers, and the Pune firewall is configured for asymmetric return traffic.

Client VPN. AWS recommends a client CIDR with twice the addresses needed for peak concurrent connections, because part of the range supports the endpoint's availability model. Peak concurrency is 1,500, so 3,000 addresses are needed; a /22 has only 1,024, so the design uses a /20 with 4,096, from an unused private range, 10.250.0.0/20, that collides with no VPC, no site and no typical home network. Concurrent connection capacity depends on associations: 7,000 with one subnet and 36,500 with two, so two associations in different Availability Zones give both headroom and resilience.

Reachability. The endpoint lives in a small hub VPC attached to the Transit Gateway. Endpoint routes send the six workload CIDRs and both site ranges into the hub subnets, and the hub VPC's route table sends them to the Transit Gateway. Authorization rules grant engineering the workload VPCs, finance one application VPC, and nobody the Pune management network. Split tunnel is on, so video calls do not hairpin through AWS; each user connection has a baseline of 50 Mbps, which suits internal traffic but would be wasted on everything else. The Transit Gateway route tables need return routes for 10.250.0.0/20 to the hub VPC, or replies from on premises have nowhere to go.

Operations and monitoring

  • Alarm on each tunnel's state in CloudWatch, per tunnel, not per connection: one tunnel down is the warning that precedes an outage during the next maintenance.
  • Log BGP session changes on the customer gateway and correlate them with AWS tunnel events.
  • Enable Client VPN connection logging to CloudWatch Logs; it is the only record of which user held which client address.
  • Revoke lost or leaver certificates through the revocation list, which holds up to 20,000 entries, and remember the change can take hours to apply.
  • Re-check the client CIDR's headroom as the workforce grows; the only fix for an exhausted range is a new endpoint.

Failure modes and trade-offs

  • One tunnel configured. Works until AWS maintains that tunnel. Configure both.
  • Static routing with multiple tunnels. No ECMP and slow failover. Use BGP.
  • Too many prefixes. The 100-route limit on a virtual private gateway cannot be raised. Summarise or move to a Transit Gateway.
  • MSS not clamped. Large transfers hang while small ones work.
  • Overlapping CIDRs. A client CIDR or home LAN that overlaps a VPC or office range makes those destinations unreachable. Client VPN also expects client LANs in private ranges and forces all LAN traffic into the tunnel otherwise.
  • Routes without authorization. Users connect and reach nothing.
  • Assuming a VPN is a firewall. Authorization rules work at CIDR level. Least privilege inside the VPC still needs security groups and identity-aware access; see zero-trust designs and, for self-managed tunnels, WireGuard as an alternative.

The overall trade-off is operations against control. Managed VPN removes servers to patch and scale and gives you two tunnels in different Availability Zones for free, at the cost of hard limits, slow configuration changes and internet-quality latency. Direct Connect adds predictable bandwidth at higher cost and lead time; a self-managed VPN on EC2 gives full control and all the operational work that comes with it.

What to do next

  1. Inventory every on-premises prefix and remote-user population, and choose non-overlapping client CIDRs before creating anything.
  2. Terminate Site-to-Site connections on a Transit Gateway with BGP, both tunnels configured, and ECMP or Large Bandwidth Tunnels where bandwidth needs it.
  3. Clamp MSS on every customer gateway and test with a large transfer, not a ping.
  4. Create the Client VPN endpoint with SAML, a client CIDR sized at twice peak concurrency, and associations in at least two Availability Zones.
  5. Write authorization rules per group and reference the endpoint's security group from target security groups.
  6. Alarm on per-tunnel state, enable connection logging, and rehearse the Direct Connect-to-VPN failover.
Key takeaway: Site-to-Site VPN joins networks with pairs of IPsec tunnels; Client VPN joins people to a VPC. Terminate site connections on a Transit Gateway with BGP and both tunnels configured, respect the fixed route limits, clamp MSS to fit the 1,446-byte MTU, and use ECMP or 5 Gbps Large Bandwidth Tunnels when one standard tunnel is not enough. For Client VPN, choose a non-overlapping client CIDR of at least twice peak concurrency before creation, associate subnets in several Availability Zones, authenticate with SAML, and remember that routes and authorization rules must both allow a flow.