AWS Direct Connect is a private network path between your network and AWS that does not cross the public internet. You get a physical Ethernet port, or a slice of a partner's, at a colocation facility, and run BGP sessions over VLANs on it. Teams adopt it for predictable latency, high sustained throughput, lower data-transfer-out rates than internet egress, and because some regulators and security teams want private links for hybrid workloads.
It is also a real piece of network engineering. A Direct Connect link is a single fiber into a single device until you design otherwise; it is unencrypted by default; and its routing is plain BGP with AWS-specific quotas and community tags. This article explains the parts from the fiber up, shows how traffic is chosen between paths, walks through a two-site design with the CLI, and lists the failures that take hybrid networks down. It assumes you know VPC routing.
The layers: connection, virtual interface, gateway
There are three layers, and most confusion comes from mixing them up.
The connection is layer 1 and 2: a port on an AWS device at a Direct Connect location. A dedicated connection is a whole port at 1, 10, 100 or 400 Gbps, and you arrange the cross-connect in the facility with the Letter of Authorization AWS issues. A hosted connection is provisioned by a Direct Connect Partner on its own port, at speeds from 50 Mbps up to 25 Gbps. The difference matters later: a dedicated connection can carry up to 51 virtual interfaces, while a hosted connection carries exactly one.
A virtual interface, or VIF, is layer 3: a VLAN tag, a pair of peer IP addresses, a BGP session and a type. A private VIF reaches VPCs through a virtual private gateway or a Direct Connect gateway. A transit VIF reaches Transit Gateways or AWS Cloud WAN through a Direct Connect gateway. A public VIF reaches AWS public service endpoints such as S3 using public IP addresses you own.
The Direct Connect gateway is a global object that sits between VIFs and the gateways in your VPC side. It lets a VIF at a location associated with one Region reach VPCs in other Regions. It is not a router between your sites: traffic entering from one VIF is not sent back out another, unless you enable SiteLink on the VIFs, and it does not route between the VPCs attached to it.
Choosing the VIF type
| Type | Reaches | Scale | Use it when |
|---|---|---|---|
| Private VIF | One VGW directly, or up to 20 VGWs via a DX gateway | Up to 50 public/private VIFs per dedicated connection | Few VPCs, simple topology |
| Transit VIF | Up to 6 Transit Gateways via a DX gateway | Up to 4 per dedicated connection | Many VPCs already on Transit Gateway |
| Public VIF | AWS public prefixes in all public Regions by default | 1,000 routes from you per session | Bulk transfer to S3 or other public endpoints, or IPsec over a public VIF |
With more than a handful of VPCs, use a transit VIF into a Direct Connect gateway associated with a Transit Gateway per Region: you pay Transit Gateway data processing but get one routing domain and far fewer BGP sessions. A private VIF to a VGW stays the simplest pattern for one or two VPCs.
The quotas that shape designs
Direct Connect has hard limits that turn into architecture constraints. The ones that bite are listed here, taken from the service quotas page; re-check them when you design, because AWS adjusts them over time.
- Routes you advertise on a private or transit VIF: 100 each for IPv4 and IPv6 by default, raisable to 1,000 with prefix controls. Advertise more than your allocation and the BGP session goes idle, which means the VIF is down.
- Routes you advertise on a public VIF: 1,000, not raisable.
- Prefixes AWS advertises to you per Transit Gateway on a transit VIF: 200 combined IPv4 and IPv6. These come from the allowed prefixes list on the gateway association, not from VPC CIDRs automatically.
- Per Direct Connect gateway: 20 virtual private gateways, 6 Transit Gateways and 30 private or transit VIFs.
- Link aggregation groups: up to 4 connections below 100 Gbps, or 2 at 100 Gbps, all at the same speed and location.
The first limit makes summarization mandatory: a network that leaks hundreds of /24s into BGP will one day exceed it and drop the session. Filter outbound advertisements to a few aggregates and alert well below the limit.
How AWS picks a path
Once you have more than one VIF, BGP attributes decide which path traffic from AWS to your network takes. For private and transit VIFs, AWS evaluates the longest prefix first. For equal prefixes it uses local preference, which you control with community tags on the routes you advertise: 7224:7100 low, 7224:7200 medium and 7224:7300 high. Without tags, AWS prefers Direct Connect locations associated with the Region sending the traffic. Next comes AS_PATH length, which you can lengthen with prepending, then MED, which AWS advises against relying on. Paths that tie on everything are load-balanced with ECMP.
Traffic toward AWS is decided by your routers. Mirror the choice with local preference on your side, or requests leave through one location and responses return through another, and stateful firewalls drop them. ECMP explains why hashing across paths is fine for stateless routers and dangerous for stateful middleboxes.
Public VIFs have their own tags. You can limit how far AWS propagates your prefixes with 7224:9100 for the local Region, 7224:9200 for the continent and 7224:9300 for all public Regions, which is the default. AWS tags the routes it sends you with 7224:8100 for the same Region and 7224:8200 for the same continent, so you can filter to nearby services. All public VIF routes carry the NO_EXPORT community; never re-advertise them to the internet.
When Direct Connect and a Site-to-Site VPN advertise the same prefix to the same gateway, AWS prefers the Direct Connect route. That makes a VPN a natural backup, but only if you keep prefixes identical; advertise a more specific prefix over the VPN and longest-prefix match silently makes the backup primary.
Resilience: design for the fiber cut
A single connection has several single points of failure: the cross-connect, the AWS device, the location's power and the partner's circuit. AWS's Resiliency Toolkit describes three models. Development and test uses a single location. High resiliency uses one connection at each of two locations. Maximum resiliency uses two connections on separate devices at each of two locations. Production hybrid workloads should start at high resiliency and usually aim for maximum, and the underlying circuits should use diverse physical paths, which you have to ask your carrier to prove.
Failure detection is BGP's job, and BGP is slow by default. The Direct Connect hold timer defaults to 90 seconds, so a silently failed path can black-hole traffic for a minute and a half. Enable asynchronous BFD on your routers; AWS enables it on its side with a minimum interval of 300 ms and a multiplier of 3, which gives detection in about 900 ms. Then test: the Direct Connect failover testing feature can bring down BGP on a chosen VIF for a set time so you can watch traffic move, and you should do this before production depends on the link, not during an incident.
MTU, encryption and performance
A private VIF supports an MTU of 1500 or 9001 bytes; a transit VIF supports 1500 or 8500. Jumbo frames help large transfers, but every hop must agree. If the VIF is 9001 while a firewall in your data center is 1500 and drops ICMP, path MTU discovery fails and large packets vanish while pings and small requests succeed, which is the classic symptom described in path MTU discovery. Changing a VIF to jumbo can update the underlying connection and interrupt every VIF on it for up to 30 seconds, so schedule it.
Direct Connect is private, not encrypted. For encryption at layer 2, MACsec is available on 10, 100 and 400 Gbps dedicated connections at select locations and requires compatible router hardware. The alternative is IPsec: a Site-to-Site VPN over a public VIF, or a private IP VPN terminated on a Transit Gateway over a transit VIF. IPsec is available everywhere but costs throughput per tunnel and adds overhead to the MTU, so plan for several tunnels with ECMP if you need bandwidth. TLS inside the application is still the right default either way.
Watch throughput in CloudWatch: ConnectionBpsEgress, ConnectionBpsIngress and ConnectionState per connection, per-VIF bandwidth metrics, and optical light levels such as ConnectionLightLevelRx on dedicated connections. A slowly falling receive light level often precedes a hard failure.
Worked example: two data centers, two Regions
A company has data centers in Frankfurt and Dublin and VPCs on Transit Gateways in eu-central-1 and eu-west-1. On-premises uses 10.0.0.0/12; AWS uses 10.64.0.0/12. The design: one dedicated 10 Gbps connection from each data center to a nearby location, each with a transit VIF into one Direct Connect gateway associated with both Transit Gateways. Frankfurt is the primary path and Dublin the backup.
# 1. A Direct Connect gateway with its own private ASN
aws directconnect create-direct-connect-gateway \
--direct-connect-gateway-name dx-eu --amazon-side-asn 64600
# 2. Associate each Transit Gateway, listing what AWS may advertise on-prem
aws directconnect create-direct-connect-gateway-association \
--direct-connect-gateway-id <dxgw-id> --gateway-id <tgw-euc1-id> \
--add-allowed-prefixes-to-direct-connect-gateway cidr=10.64.0.0/13
aws directconnect create-direct-connect-gateway-association \
--direct-connect-gateway-id <dxgw-id> --gateway-id <tgw-euw1-id> \
--add-allowed-prefixes-to-direct-connect-gateway cidr=10.72.0.0/13
# 3. A transit VIF on each connection
aws directconnect create-transit-virtual-interface \
--connection-id <dxcon-fra> \
--new-transit-virtual-interface '{"virtualInterfaceName": "fra-transit",
"vlan": 101, "asn": 65010, "mtu": 8500, "addressFamily": "ipv4",
"directConnectGatewayId": "<dxgw-id>"}'On the Frankfurt router, advertise the aggregate with a high-preference community for the primary role and let Dublin advertise the same aggregate at low preference. The fragment below uses FRR syntax; adapt it to your platform.
router bgp 65010
neighbor 169.254.10.1 remote-as 64600
neighbor 169.254.10.1 password <md5-key>
neighbor 169.254.10.1 bfd
address-family ipv4 unicast
network 10.0.0.0/12
neighbor 169.254.10.1 route-map TO-AWS out
neighbor 169.254.10.1 prefix-list FROM-AWS in
!
ip prefix-list ONPREM seq 10 permit 10.0.0.0/12
ip prefix-list FROM-AWS seq 10 permit 10.64.0.0/12 le 13
route-map TO-AWS permit 10
match ip address prefix-list ONPREM
set community 7224:7300The design advertises one prefix, so the route quota is irrelevant, and uses communities rather than MED. To make each Region prefer its local data center rather than one global primary, advertise more specific halves of 10.0.0.0/12 from each site with different communities, and keep the aggregate on both as the fallback. Finally, run a failover test on each VIF and confirm with flow logs and traceroutes that traffic moves within about a second and returns symmetrically.
Failure modes
| Symptom | Cause | Prevention |
|---|---|---|
| VIF down after an on-prem change | More prefixes than the allocation; BGP session idle | Outbound prefix filter, aggregation, alerting |
| Large transfers hang, small ones work | MTU mismatch with ICMP blocked | Consistent MTU end to end, allow ICMP fragmentation-needed |
| Firewall drops return traffic | Asymmetric routing across two locations | Match local preference in both directions |
| Backup VPN carries all traffic | VPN advertises more specific prefixes | Identical prefixes on both paths |
| Minutes of outage on a fiber cut | No BFD, 90 s hold timer | Enable BFD; test failover |
| Both 'redundant' links fail together | Shared conduit, device or location | Diverse locations and verified physical diversity |
| VPC CIDR not reachable from on-prem | Missing from allowed prefixes on the association | Update the allowed prefix list |
For the wider picture of hybrid and multi-cloud connectivity patterns, including when a VPN alone is enough, see cloud networking.
Cost and trade-offs
You pay per port-hour for the connection, or the partner's price for a hosted connection, plus data transfer out of AWS at Direct Connect rates, which are lower than internet egress in most Regions. Inbound transfer is not charged. With a transit VIF you also pay Transit Gateway attachment and data processing. The break-even against VPN over the internet depends on volume; below a few terabytes a month a VPN is often cheaper and faster to provision, while above that the lower transfer rate and consistent performance usually justify Direct Connect. Provisioning time is the hidden cost: cross-connects and carrier circuits take weeks, so order the second path when you order the first.
What to do next
- Write down your required availability and pick a Resiliency Toolkit model to match; do not run production on a single connection.
- Choose private or transit VIFs based on VPC count, and create one Direct Connect gateway per routing domain.
- Summarize on-premises routes to a handful of aggregates and add an outbound prefix filter with an alert below the quota.
- Set local preference communities for primary and backup paths and mirror the preference on your own routers.
- Enable BFD, then run a failover test on every VIF and record how long traffic takes to move.
- Decide on encryption: MACsec where supported, otherwise IPsec over the link, plus TLS in applications.
- Pick an MTU per VIF, verify it end to end, and allow the ICMP messages path MTU discovery needs.
- Add CloudWatch alarms on connection state, BGP status, bandwidth and light levels.