AWS Direct Connect is a private network path between your network and AWS that does not cross the public internet. You get a physical Ethernet port, or a slice of a partner's, at a colocation facility, and run BGP sessions over VLANs on it. Teams adopt it for predictable latency, high sustained throughput, lower data-transfer-out rates than internet egress, and because some regulators and security teams want private links for hybrid workloads.

It is also a real piece of network engineering. A Direct Connect link is a single fiber into a single device until you design otherwise; it is unencrypted by default; and its routing is plain BGP with AWS-specific quotas and community tags. This article explains the parts from the fiber up, shows how traffic is chosen between paths, walks through a two-site design with the CLI, and lists the failures that take hybrid networks down. It assumes you know VPC routing.

Advertisement

The layers: connection, virtual interface, gateway

Direct Connect: physical port, VLAN virtual interfaces, gatewaysOn-prem routerBGP, your ASNDX locationcross-connect to AWS devicefiberPrivate VIFVLAN + BGPTransit VIFVLAN + BGPPublic VIFVLAN + BGPDX gatewayglobal objectTransit gatewayper RegionAWS public endpointsS3, DynamoDB ...VPC A (VGW)Region 1VPCs B..Nvia TGW attachmentsprivate VIF pathNo encryption by default: add MACsec or IPsec
A physical connection at a Direct Connect location carries 802.1Q VLANs. Each VLAN is a virtual interface with its own BGP session, and the interface type decides what it can reach.

There are three layers, and most confusion comes from mixing them up.

The connection is layer 1 and 2: a port on an AWS device at a Direct Connect location. A dedicated connection is a whole port at 1, 10, 100 or 400 Gbps, and you arrange the cross-connect in the facility with the Letter of Authorization AWS issues. A hosted connection is provisioned by a Direct Connect Partner on its own port, at speeds from 50 Mbps up to 25 Gbps. The difference matters later: a dedicated connection can carry up to 51 virtual interfaces, while a hosted connection carries exactly one.

A virtual interface, or VIF, is layer 3: a VLAN tag, a pair of peer IP addresses, a BGP session and a type. A private VIF reaches VPCs through a virtual private gateway or a Direct Connect gateway. A transit VIF reaches Transit Gateways or AWS Cloud WAN through a Direct Connect gateway. A public VIF reaches AWS public service endpoints such as S3 using public IP addresses you own.

The Direct Connect gateway is a global object that sits between VIFs and the gateways in your VPC side. It lets a VIF at a location associated with one Region reach VPCs in other Regions. It is not a router between your sites: traffic entering from one VIF is not sent back out another, unless you enable SiteLink on the VIFs, and it does not route between the VPCs attached to it.

Choosing the VIF type

TypeReachesScaleUse it when
Private VIFOne VGW directly, or up to 20 VGWs via a DX gatewayUp to 50 public/private VIFs per dedicated connectionFew VPCs, simple topology
Transit VIFUp to 6 Transit Gateways via a DX gatewayUp to 4 per dedicated connectionMany VPCs already on Transit Gateway
Public VIFAWS public prefixes in all public Regions by default1,000 routes from you per sessionBulk transfer to S3 or other public endpoints, or IPsec over a public VIF

With more than a handful of VPCs, use a transit VIF into a Direct Connect gateway associated with a Transit Gateway per Region: you pay Transit Gateway data processing but get one routing domain and far fewer BGP sessions. A private VIF to a VGW stays the simplest pattern for one or two VPCs.

Advertisement

The quotas that shape designs

Direct Connect has hard limits that turn into architecture constraints. The ones that bite are listed here, taken from the service quotas page; re-check them when you design, because AWS adjusts them over time.

  • Routes you advertise on a private or transit VIF: 100 each for IPv4 and IPv6 by default, raisable to 1,000 with prefix controls. Advertise more than your allocation and the BGP session goes idle, which means the VIF is down.
  • Routes you advertise on a public VIF: 1,000, not raisable.
  • Prefixes AWS advertises to you per Transit Gateway on a transit VIF: 200 combined IPv4 and IPv6. These come from the allowed prefixes list on the gateway association, not from VPC CIDRs automatically.
  • Per Direct Connect gateway: 20 virtual private gateways, 6 Transit Gateways and 30 private or transit VIFs.
  • Link aggregation groups: up to 4 connections below 100 Gbps, or 2 at 100 Gbps, all at the same speed and location.

The first limit makes summarization mandatory: a network that leaks hundreds of /24s into BGP will one day exceed it and drop the session. Filter outbound advertisements to a few aggregates and alert well below the limit.

How AWS picks a path

Once you have more than one VIF, BGP attributes decide which path traffic from AWS to your network takes. For private and transit VIFs, AWS evaluates the longest prefix first. For equal prefixes it uses local preference, which you control with community tags on the routes you advertise: 7224:7100 low, 7224:7200 medium and 7224:7300 high. Without tags, AWS prefers Direct Connect locations associated with the Region sending the traffic. Next comes AS_PATH length, which you can lengthen with prepending, then MED, which AWS advises against relying on. Paths that tie on everything are load-balanced with ECMP.

Traffic toward AWS is decided by your routers. Mirror the choice with local preference on your side, or requests leave through one location and responses return through another, and stateful firewalls drop them. ECMP explains why hashing across paths is fine for stateless routers and dangerous for stateful middleboxes.

Public VIFs have their own tags. You can limit how far AWS propagates your prefixes with 7224:9100 for the local Region, 7224:9200 for the continent and 7224:9300 for all public Regions, which is the default. AWS tags the routes it sends you with 7224:8100 for the same Region and 7224:8200 for the same continent, so you can filter to nearby services. All public VIF routes carry the NO_EXPORT community; never re-advertise them to the internet.

When Direct Connect and a Site-to-Site VPN advertise the same prefix to the same gateway, AWS prefers the Direct Connect route. That makes a VPN a natural backup, but only if you keep prefixes identical; advertise a more specific prefix over the VPN and longest-prefix match silently makes the backup primary.

Resilience: design for the fiber cut

Maximum resiliency: two locations, two devices each, one BGP session per pathData center 1routers R1, R2Data center 2routers R3, R4Location A dev 1Location A dev 2Location B dev 1Location B dev 2DX gateway4 transit VIFs attached7224:7300 on primary pair7224:7100 on backup pair
The maximum resiliency model: separate connections on separate devices at two locations, preferably from two data centers. Community tags choose the active pair; BFD detects failure in under a second.

A single connection has several single points of failure: the cross-connect, the AWS device, the location's power and the partner's circuit. AWS's Resiliency Toolkit describes three models. Development and test uses a single location. High resiliency uses one connection at each of two locations. Maximum resiliency uses two connections on separate devices at each of two locations. Production hybrid workloads should start at high resiliency and usually aim for maximum, and the underlying circuits should use diverse physical paths, which you have to ask your carrier to prove.

Failure detection is BGP's job, and BGP is slow by default. The Direct Connect hold timer defaults to 90 seconds, so a silently failed path can black-hole traffic for a minute and a half. Enable asynchronous BFD on your routers; AWS enables it on its side with a minimum interval of 300 ms and a multiplier of 3, which gives detection in about 900 ms. Then test: the Direct Connect failover testing feature can bring down BGP on a chosen VIF for a set time so you can watch traffic move, and you should do this before production depends on the link, not during an incident.

MTU, encryption and performance

A private VIF supports an MTU of 1500 or 9001 bytes; a transit VIF supports 1500 or 8500. Jumbo frames help large transfers, but every hop must agree. If the VIF is 9001 while a firewall in your data center is 1500 and drops ICMP, path MTU discovery fails and large packets vanish while pings and small requests succeed, which is the classic symptom described in path MTU discovery. Changing a VIF to jumbo can update the underlying connection and interrupt every VIF on it for up to 30 seconds, so schedule it.

Direct Connect is private, not encrypted. For encryption at layer 2, MACsec is available on 10, 100 and 400 Gbps dedicated connections at select locations and requires compatible router hardware. The alternative is IPsec: a Site-to-Site VPN over a public VIF, or a private IP VPN terminated on a Transit Gateway over a transit VIF. IPsec is available everywhere but costs throughput per tunnel and adds overhead to the MTU, so plan for several tunnels with ECMP if you need bandwidth. TLS inside the application is still the right default either way.

Watch throughput in CloudWatch: ConnectionBpsEgress, ConnectionBpsIngress and ConnectionState per connection, per-VIF bandwidth metrics, and optical light levels such as ConnectionLightLevelRx on dedicated connections. A slowly falling receive light level often precedes a hard failure.

Worked example: two data centers, two Regions

A company has data centers in Frankfurt and Dublin and VPCs on Transit Gateways in eu-central-1 and eu-west-1. On-premises uses 10.0.0.0/12; AWS uses 10.64.0.0/12. The design: one dedicated 10 Gbps connection from each data center to a nearby location, each with a transit VIF into one Direct Connect gateway associated with both Transit Gateways. Frankfurt is the primary path and Dublin the backup.

# 1. A Direct Connect gateway with its own private ASN
aws directconnect create-direct-connect-gateway \
  --direct-connect-gateway-name dx-eu --amazon-side-asn 64600

# 2. Associate each Transit Gateway, listing what AWS may advertise on-prem
aws directconnect create-direct-connect-gateway-association \
  --direct-connect-gateway-id <dxgw-id> --gateway-id <tgw-euc1-id> \
  --add-allowed-prefixes-to-direct-connect-gateway cidr=10.64.0.0/13

aws directconnect create-direct-connect-gateway-association \
  --direct-connect-gateway-id <dxgw-id> --gateway-id <tgw-euw1-id> \
  --add-allowed-prefixes-to-direct-connect-gateway cidr=10.72.0.0/13

# 3. A transit VIF on each connection
aws directconnect create-transit-virtual-interface \
  --connection-id <dxcon-fra> \
  --new-transit-virtual-interface '{"virtualInterfaceName": "fra-transit",
      "vlan": 101, "asn": 65010, "mtu": 8500, "addressFamily": "ipv4",
      "directConnectGatewayId": "<dxgw-id>"}'

On the Frankfurt router, advertise the aggregate with a high-preference community for the primary role and let Dublin advertise the same aggregate at low preference. The fragment below uses FRR syntax; adapt it to your platform.

router bgp 65010
 neighbor 169.254.10.1 remote-as 64600
 neighbor 169.254.10.1 password <md5-key>
 neighbor 169.254.10.1 bfd
 address-family ipv4 unicast
  network 10.0.0.0/12
  neighbor 169.254.10.1 route-map TO-AWS out
  neighbor 169.254.10.1 prefix-list FROM-AWS in
!
ip prefix-list ONPREM seq 10 permit 10.0.0.0/12
ip prefix-list FROM-AWS seq 10 permit 10.64.0.0/12 le 13
route-map TO-AWS permit 10
 match ip address prefix-list ONPREM
 set community 7224:7300

The design advertises one prefix, so the route quota is irrelevant, and uses communities rather than MED. To make each Region prefer its local data center rather than one global primary, advertise more specific halves of 10.0.0.0/12 from each site with different communities, and keep the aggregate on both as the fallback. Finally, run a failover test on each VIF and confirm with flow logs and traceroutes that traffic moves within about a second and returns symmetrically.

Failure modes

SymptomCausePrevention
VIF down after an on-prem changeMore prefixes than the allocation; BGP session idleOutbound prefix filter, aggregation, alerting
Large transfers hang, small ones workMTU mismatch with ICMP blockedConsistent MTU end to end, allow ICMP fragmentation-needed
Firewall drops return trafficAsymmetric routing across two locationsMatch local preference in both directions
Backup VPN carries all trafficVPN advertises more specific prefixesIdentical prefixes on both paths
Minutes of outage on a fiber cutNo BFD, 90 s hold timerEnable BFD; test failover
Both 'redundant' links fail togetherShared conduit, device or locationDiverse locations and verified physical diversity
VPC CIDR not reachable from on-premMissing from allowed prefixes on the associationUpdate the allowed prefix list

For the wider picture of hybrid and multi-cloud connectivity patterns, including when a VPN alone is enough, see cloud networking.

Cost and trade-offs

You pay per port-hour for the connection, or the partner's price for a hosted connection, plus data transfer out of AWS at Direct Connect rates, which are lower than internet egress in most Regions. Inbound transfer is not charged. With a transit VIF you also pay Transit Gateway attachment and data processing. The break-even against VPN over the internet depends on volume; below a few terabytes a month a VPN is often cheaper and faster to provision, while above that the lower transfer rate and consistent performance usually justify Direct Connect. Provisioning time is the hidden cost: cross-connects and carrier circuits take weeks, so order the second path when you order the first.

What to do next

  1. Write down your required availability and pick a Resiliency Toolkit model to match; do not run production on a single connection.
  2. Choose private or transit VIFs based on VPC count, and create one Direct Connect gateway per routing domain.
  3. Summarize on-premises routes to a handful of aggregates and add an outbound prefix filter with an alert below the quota.
  4. Set local preference communities for primary and backup paths and mirror the preference on your own routers.
  5. Enable BFD, then run a failover test on every VIF and record how long traffic takes to move.
  6. Decide on encryption: MACsec where supported, otherwise IPsec over the link, plus TLS in applications.
  7. Pick an MTU per VIF, verify it end to end, and allow the ICMP messages path MTU discovery needs.
  8. Add CloudWatch alarms on connection state, BGP status, bandwidth and light levels.
Key takeaway: Direct Connect gives you a private, predictable path into AWS, but it is BGP over VLANs on a physical port, with everything that implies. Pick hosted or dedicated connections by speed and VIF needs, prefer transit VIFs through a Direct Connect gateway when you have many VPCs, and respect the route quotas by summarizing. Control path selection with longest prefix and the 7224 local preference communities, keep routing symmetric, enable BFD and test failover. Build across two locations with diverse circuits, add MACsec or IPsec because the link is not encrypted, and verify MTU end to end before production traffic finds the mismatch for you.