The Border Gateway Protocol (BGP) is how independently run networks, called autonomous systems (ASes), tell each other which IP prefixes they can reach. It carries the Internet's routing table of about a million IPv4 prefixes, and it is also used inside data centres as the underlay routing protocol and to announce service addresses from load balancers and Kubernetes nodes.

BGP is unusual in two ways. It chooses routes by policy, not by shortest path, and it trusts what neighbours say unless you configure it not to. Both explain its biggest outages. This article builds BGP up from the session to the Internet: how peers talk, how a router stores and chooses routes, how policy and route security work, how fast it converges, and how to configure a dual-homed edge safely.

Sessions and messages

Two BGP speakers form a session over TCP port 179, configured explicitly on both sides; BGP has no neighbour discovery. A session between different ASes is external BGP (eBGP), and one inside an AS is internal BGP (iBGP). Each side moves through the finite-state machine in RFC 4271: Idle, Connect, Active (retrying the TCP connection), OpenSent, OpenConfirm and Established. A session stuck flapping between Active and Idle usually means TCP cannot connect: a wrong address, a firewall, or an MD5 or TCP-AO key mismatch.

Five message types do all the work. OPEN carries the AS number, a proposed hold time, the router ID and capabilities, such as multiprotocol support for IPv6 and VPNs, four-byte AS numbers, graceful restart and ADD-PATH. UPDATE carries withdrawn prefixes and new prefixes that share one set of path attributes. KEEPALIVE proves liveness. NOTIFICATION reports an error and closes the session. ROUTE-REFRESH asks a peer to resend its routes after you change inbound policy. The hold time is the lower of the two proposals; if no message arrives within it, the session drops and every route learned over it is withdrawn. RFC 4271 suggests 90 seconds with keepalives at one third of that, and many vendors default to 180 and 60.

BGP is incremental: after the initial exchange, peers send only changes. A session with a full Internet table exchanges about a million prefixes at start-up and then a steady stream of updates as networks elsewhere change.

Inside a speaker: RIBs and policy

Inside one BGP speaker: per-peer RIBs, policy on both sides, one best path per prefixPeer AeBGP, TCP 179Peer BiBGP / RRAdj-RIB-Inper peer, rawInbound policyfilter, RPKI, set LPLoc-RIBbest path per prefixRIB / FIBforwarding tableOutbound policyexport, prependAdj-RIB-Outper peerRPKI cacheRTR, validated ROAsinstallUPDATE / WITHDRAWRoutes enter per peer, pass inbound policy, compete in best-path selection,and only the winner is installed and, after outbound policy, advertised to peers.
The route-processing pipeline inside one BGP speaker, following the RFC 4271 RIB model.

RFC 4271 describes three logical tables. Each peer's routes land in its Adj-RIB-In. Inbound policy filters and modifies them, and every surviving path for a prefix competes in best-path selection. The winners form the Loc-RIB, which feeds the router's main routing table and forwarding table. Outbound policy then decides what each peer may see, forming an Adj-RIB-Out per peer. Only one best path per prefix is advertised unless ADD-PATH (RFC 7911) is negotiated.

Path attributes are the data that policy and selection work on. AS_PATH lists the ASes a route has crossed, both as a loop guard and as a distance measure. NEXT_HOP says where to send traffic. LOCAL_PREF expresses preference inside your AS. MULTI_EXIT_DISC (MED) is a hint to a neighbour about which of your links to prefer. ORIGIN records how the route entered BGP. COMMUNITIES (RFC 1997) and large communities (RFC 8092) are tags that carry policy meaning between networks.

Best-path selection

When several paths to one prefix survive inbound policy, the router picks one by comparing attributes in a fixed order. First it discards paths whose next hop is unreachable. Then, in the order most implementations use: highest LOCAL_PREF, locally originated routes, shortest AS_PATH, lowest ORIGIN, lowest MED (compared only between paths from the same neighbouring AS), eBGP over iBGP, lowest IGP cost to the next hop, then tie-breakers such as oldest path, lowest router ID and lowest neighbour address. Vendors add steps; Cisco and FRR, for example, put a local-only weight first. The code below captures the core order.

from dataclasses import dataclass

@dataclass
class Path:
    local_pref: int = 100
    locally_originated: bool = False
    as_path: tuple = ()
    origin: int = 0              # IGP=0, EGP=1, INCOMPLETE=2
    med: int = 0
    neighbor_as: int = 0
    ebgp: bool = True
    igp_cost: int = 0
    router_id: str = ""

def better(a, b):
    """Simplified RFC 4271-style comparison with common vendor steps (weight omitted)."""
    if a.local_pref != b.local_pref:
        return a if a.local_pref > b.local_pref else b
    if a.locally_originated != b.locally_originated:
        return a if a.locally_originated else b
    if len(a.as_path) != len(b.as_path):
        return a if len(a.as_path) < len(b.as_path) else b
    if a.origin != b.origin:
        return a if a.origin < b.origin else b
    if a.neighbor_as == b.neighbor_as and a.med != b.med:   # MED only within one neighbour AS
        return a if a.med < b.med else b
    if a.ebgp != b.ebgp:
        return a if a.ebgp else b
    if a.igp_cost != b.igp_cost:                            # hot-potato routing
        return a if a.igp_cost < b.igp_cost else b
    return a if a.router_id < b.router_id else b            # final tie-breaks vary by vendor

Two consequences are worth knowing. First, LOCAL_PREF beats AS_PATH, so your policy overrides topology; this is how you make a cheap link primary. Second, because MED is compared only within one neighbour AS, pairwise comparison can depend on the order in which paths arrived. That is why FRR offers bgp deterministic-med and Cisco an equivalent, which group paths by neighbour AS first. Enable it everywhere or nowhere.

iBGP, route reflectors and next hops

Routes learned over iBGP are not re-advertised to other iBGP peers, which prevents loops inside an AS but means every iBGP speaker must peer with every other. That is n(n-1)/2 sessions: 45 for 10 routers, 4,950 for 100. Route reflectors (RFC 4456) break the mesh: clients peer only with reflectors, and reflectors re-advertise to clients, using ORIGINATOR_ID and CLUSTER_LIST to prevent loops. Run at least two reflectors per cluster.

Two details bite. eBGP next hops are carried unchanged into iBGP, so internal routers must reach the external peer's address, or the edge sets next-hop-self. And a reflector advertises only its own best path, hiding alternatives from clients; ADD-PATH restores them, so clients can fail over or load-share, as described in ECMP routing.

Policy: what you accept, prefer and announce

Inbound policy decides what you accept and how much you prefer it. Outbound policy decides what you announce. RFC 8212 requires eBGP speakers to accept and announce nothing without explicit policy, and FRR follows it with bgp ebgp-requires-policy on by default in its traditional profile.

To steer outbound traffic, set LOCAL_PREF on routes from the preferred provider. Steering inbound traffic is harder because other networks decide. You can prepend your AS to make a path look longer, announce more-specific prefixes on the preferred link, or use the provider's documented communities, for example to lower local preference inside their network. Prepending is weak against networks that prefer customer routes by LOCAL_PREF, and more-specifics add to the global table.

Filtering discipline is what keeps you from breaking the Internet. Announce only your own prefixes and your customers', from an explicit prefix list. On inbound, drop martians and bogons, prefixes longer than /24 for IPv4 or /48 for IPv6, and RPKI-invalid routes. Set a maximum-prefix limit on every session so that a peer leaking a full table resets the session instead of filling your routers' memory.

Route security: hijacks, leaks and RPKI

BGP has no built-in check that an AS may originate a prefix or that a path is real. A hijack is an origination of someone else's prefix, as when Pakistan Telecom announced part of YouTube's space in 2008. A route leak is a legitimate route propagated beyond its intended scope, as in 2019 when a customer leaked more-specifics learned from one provider to another, which accepted and propagated them.

Defences are layered. Prefix filters built from Internet Routing Registry data stop customers announcing space they do not hold. RPKI route origin validation (RFC 6811) checks each route against cryptographically signed Route Origin Authorizations: a route is valid, invalid (wrong origin AS or too specific), or not found. Routers get validated data over the RTR protocol from a local validator such as Routinator or rpki-client, and drop invalids. Publish ROAs for your own prefixes so others can do the same. ROV does not detect a forged path with a valid origin. BGP roles (RFC 9234) let peers declare customer, provider or peer relationships and mark routes with an Only-to-Customer attribute to stop leaks, and ASPA, which validates AS_PATH against signed provider lists, was still an IETF Internet-Draft, not an RFC, at the time of writing. BGPsec (RFC 8205) protects whole paths but has very little deployment.

Convergence: detection, propagation and restart

Convergence time is failure detection plus propagation. Detection by hold timer is slow: with a 90-second hold time a silent failure, such as a dead optic behind a switch, can take up to 90 seconds to notice. Bidirectional Forwarding Detection (BFD, RFC 5880) runs a fast hello protocol under the session; three missed 300 ms packets detect failure in under a second and tear the session down.

Propagation is slowed by the Minimum Route Advertisement Interval (MRAI), which batches updates to a peer; RFC 4271 suggests 30 seconds for eBGP and 5 for iBGP, but many implementations use lower defaults, so check yours. When a prefix disappears, routers may try successively longer backup paths (path hunting) before withdrawing it, sending many updates. Graceful restart (RFC 4724) lets a peer keep forwarding on stale routes while its control plane restarts, which helps planned upgrades but can hide real failures, so pair it with BFD. Route flap damping (RFC 2439) suppresses unstable prefixes; RFC 7196 recommends much higher thresholds than the original defaults.

Worked example: a dual-homed edge

An enterprise in AS 64500 holds 198.51.100.0/24 and buys transit from AS 64510 (primary) and AS 64520 (backup). It needs to receive full tables, prefer transit A for inbound traffic, drop invalids and fail over within a second. The FRR configuration below expresses that.

router bgp 64500
 bgp router-id 192.0.2.1
 neighbor 203.0.113.1 remote-as 64510
 neighbor 203.0.113.1 description transit-a
 neighbor 203.0.113.1 password use-a-real-secret
 neighbor 203.0.113.1 bfd
 neighbor 198.18.0.1 remote-as 64520
 neighbor 198.18.0.1 description transit-b
 neighbor 198.18.0.1 bfd
 address-family ipv4 unicast
  network 198.51.100.0/24
  neighbor 203.0.113.1 route-map TRANSIT-IN in
  neighbor 203.0.113.1 route-map EXPORT-A out
  neighbor 203.0.113.1 maximum-prefix 1200000 90
  neighbor 198.18.0.1 route-map TRANSIT-IN in
  neighbor 198.18.0.1 route-map EXPORT-B out
  neighbor 198.18.0.1 maximum-prefix 1200000 90
 exit-address-family
!
ip prefix-list OURS seq 10 permit 198.51.100.0/24
ip prefix-list MARTIANS seq 10 permit 10.0.0.0/8 le 32
ip prefix-list MARTIANS seq 20 permit 192.168.0.0/16 le 32
ip prefix-list MARTIANS seq 30 permit 172.16.0.0/12 le 32
ip prefix-list MARTIANS seq 40 permit 100.64.0.0/10 le 32
ip prefix-list MARTIANS seq 50 permit 127.0.0.0/8 le 32
ip prefix-list MARTIANS seq 60 permit 169.254.0.0/16 le 32
ip prefix-list MARTIANS seq 70 permit 0.0.0.0/8 le 32
ip prefix-list MARTIANS seq 80 permit 224.0.0.0/3 le 32
ip prefix-list MARTIANS seq 90 permit 0.0.0.0/0 ge 25
! abbreviated: also add the documentation and benchmarking ranges and any unallocated space
!
rpki
 rpki cache 192.0.2.50 3323 preference 1
!
route-map TRANSIT-IN deny 10
 match rpki invalid
route-map TRANSIT-IN deny 20
 match ip address prefix-list MARTIANS
route-map TRANSIT-IN permit 30
!
route-map EXPORT-A permit 10
 match ip address prefix-list OURS
route-map EXPORT-B permit 10
 match ip address prefix-list OURS
 set as-path prepend 64500 64500

Outbound traffic uses AS path length and then IGP cost to pick between the two transits per destination; to force A, set a higher LOCAL_PREF in a separate inbound route-map for it. Inbound, the double prepend towards B makes that path look longer, so most networks prefer A; confirm with a looking glass or public route collectors. maximum-prefix 1200000 90 warns at 90 percent and resets the session above 1.2 million prefixes. The network statement announces the prefix only if it exists in the routing table, so add a blackhole static route for the aggregate. RPKI matching needs bgpd started with the rpki module, neighbor ... bfd needs the bfdd daemon running, and the rpki cache syntax differs between FRR releases, so check yours. Failure of transit A's link drops BFD within a second, both sessions reconverge, and inbound traffic shifts to B as the withdrawal propagates, typically in seconds to tens of seconds. For announcing one prefix from many sites, see anycast; for the cloud equivalent of this edge, see AWS Direct Connect.

Failure modes and trade-offs

  • Missing export filter. Redistributing a full table to a provider is the classic leak. Use explicit prefix lists and max-prefix on both sides.
  • Withdrawing yourself off the Internet. In 2021 Facebook's backbone maintenance cut off its data centres, and its DNS servers then withdrew their prefixes. Keep out-of-band access that does not depend on the network you are changing.
  • Silent failures. Without BFD a black-holed link can keep a session up until the hold timer expires.
  • iBGP next-hop unreachable. Routes sit in the table but are not used; set next-hop-self or carry the links in the IGP.
  • Asymmetric routing. Inbound and outbound paths differ, which breaks stateful firewalls that see only one direction.
  • Trade-off: more specific announcements and aggressive timers buy control and speed at the cost of global table growth and update churn.

What to do next

  1. Inventory every BGP session with its peer AS, relationship, filters, max-prefix limit and BFD status.
  2. Replace any permit-all policy with explicit inbound and outbound route-maps, and confirm RFC 8212 behaviour is on.
  3. Create ROAs for all prefixes you originate, run two RPKI validators, and drop invalids on external sessions.
  4. Enable BFD on eBGP sessions where the peer supports it, and review graceful restart settings with it.
  5. Test inbound traffic engineering from outside, using looking glasses and public route collectors, before relying on it.
  6. Monitor your prefixes with BMP (RFC 7854) or an external service, and alert on unexpected origins or paths.
Key takeaway: BGP is a policy-driven path-vector protocol that believes its neighbours. Each speaker keeps routes per peer, filters them, chooses one best path per prefix by an ordered list of attributes, and announces only what outbound policy allows. Safety comes from explicit filters, maximum-prefix limits, RPKI origin validation and published ROAs. Speed comes from BFD and sensible timers. Design your policy on paper first, then verify it from outside your network.