Most application developers never configure a router, yet BGP decides whether their users can reach them. When a status page says "connectivity issues for some customers in one region", or a payment provider is reachable from your laptop but not from your servers, the cause is often routing: a path withdrawn, a route leaked, a prefix announced by the wrong network. Those failures do not show up in your logs because the packets never arrive.

This article explains BGP as far as an application developer needs it: how your service's addresses become reachable across the internet, why routing incidents are almost always partial, how to look up what the internet currently believes about your prefixes, how convergence time interacts with your timeouts and retries, and where BGP appears inside your own infrastructure. For router internals such as RIBs and best-path selection, see BGP architecture, in depth.

Advertisement

The model in five minutes

The internet is a network of more than 70,000 independently run networks called autonomous systems (ASes), each identified by a number such as AS 64500. Your cloud provider is one; so is each ISP your users connect through. Addresses are allocated in blocks called prefixes, such as 203.0.113.0/24.

BGP is the protocol ASes use to tell each other which prefixes they can reach. Your provider's AS announces your prefix to its neighbours. Each neighbour adds its own AS number to the AS path and passes the announcement on, subject to its policy. An ISP in Singapore may end up with several candidate paths to your prefix; it picks one according to its own preferences, often based on commercial relationships before path length, and forwards your users' packets along it. When a link fails, the AS that lost it sends a withdrawal, and others fall back to their next-best path, if they have one.

Three properties matter for developers. BGP carries reachability, not performance: it does not know a path is congested or slow. Each network chooses independently, so there is no global view and no single place that decides how users reach you. And the most specific prefix wins: if anyone announces 203.0.113.0/25, routers send traffic for addresses in that half there, whatever the /24 says. That last rule is behind most hijacks.

Your prefix reaches users only as far as the announcements of other networks carry itYour service203.0.113.0/24Origin AS 64500cloud or your edgeTransit AAS 65010Transit BAS 65020IXP peerAS 65030announceISP Europepath AISP Asiapath BISP USpath via IXPEach ISP picks one best path by its own policy; a withdrawal on one link changes some users' paths, not allTraffic flows the opposite way to the announcements: users follow the paths they learned toward your prefix
Announcements flow outward from the origin AS; each network keeps the best path by its own policy, so different users reach the same service over different routes.

Why routing outages are partial

Because every network chooses its own path, a routing failure removes some paths and leaves others. Users whose ISP had chosen the failed path lose connectivity until their ISP converges on an alternative; users on other paths notice nothing. The result is the classic confusing incident: monitors green, tickets from one country, nothing reproducible from home.

Partial failures also appear between your own services. A third-party API hosted on a different provider can be unreachable from one of your regions only, because the path from that region's provider broke. Retrying from the same host takes the same path and fails the same way.

The practical lesson is to measure reachability from many networks. Monitoring from one vantage point measures one path. Real-user monitoring, probes from several providers and countries, and connection error rates split by client ASN make a routing incident visible as a pattern instead of noise.

Advertisement

Looking up what the internet believes

When you suspect routing, three questions come first: who originates my prefix, how widely is it seen, and is that origin authorised? Public route collectors such as RIPE RIS and RouteViews record what many networks see, and several free tools sit on top of them. Team Cymru's whois service maps an IP address to its announcing AS and prefix:

$ whois -h whois.cymru.com " -v 203.0.113.10"
AS      | IP            | BGP Prefix      | CC | Registry | Allocated  | AS Name
64500   | 203.0.113.10  | 203.0.113.0/24  | .. | ...      | ...        | EXAMPLE-NET

RIPEstat's data API returns visibility, origins and RPKI status as JSON, which makes it easy to put in a runbook script or a scheduled check. The script below uses only the standard library; the endpoint names are documented on stat.ripe.net, but inspect the JSON your query returns before relying on a particular field.

import json, urllib.request

def ripestat(endpoint, **params):
    query = "&".join(f"{k}={v}" for k, v in params.items())
    url = f"https://stat.ripe.net/data/{endpoint}/data.json?{query}"
    with urllib.request.urlopen(url, timeout=20) as resp:
        return json.load(resp)["data"]

prefix, origin = "203.0.113.0/24", "AS64500"   # replace with your own

status = ripestat("routing-status", resource=prefix)
print("visibility:", json.dumps(status.get("visibility"), indent=1))
print("origins seen:", [o.get("origin") for o in status.get("origins", [])])

rpki = ripestat("rpki-validation", resource=origin, prefix=prefix)
print("RPKI status:", rpki.get("status"))          # valid / invalid / unknown

Read the results as an incident checklist. Visibility well below the number of collectors means the prefix is withdrawn or filtered in places. More than one origin AS means someone else is announcing it, which is either a misconfiguration or a hijack. An RPKI status of invalid means networks that enforce route origin validation will drop the route. Looking glasses run by ISPs and exchanges then show the actual AS path from one network, which is how you find where a path breaks.

Convergence time versus your timeouts

When a path fails, recovery has two phases: detecting the failure and propagating the change. If the link goes physically down, routers notice almost immediately. If the far side stops responding without the link dropping, BGP relies on its hold timer; RFC 4271 suggests 90 seconds, and some vendors default to 180. Operators who need faster detection pair BGP with BFD, which can detect failure in well under a second. Propagation then takes from seconds to a few minutes as withdrawals ripple outward and networks try alternative paths, sometimes several in turn.

Compare those numbers with your client configuration. A request with a 30-second timeout and three retries on the same connection pool can spend 90 seconds failing against a path that is being repaired, and hold a thread for all of it. Better client behaviour assumes routing can fail briefly and partially:

  • Short connect timeouts, a few seconds, separate from read timeouts. A connect that has not finished in three seconds on a healthy internet path rarely finishes.
  • Try more than one address. If DNS returns several A or AAAA records, try the next one on connect failure, and race IPv4 and IPv6 as Happy Eyeballs (RFC 8305) does. Different addresses may sit in different prefixes with different paths.
  • Fresh connections on retry. Retrying on the same dead pooled connection waits for TCP to give up. Discard connections that saw timeouts.
  • Retry budgets and backoff. When a path returns, every client retrying at once can overload the service that was unreachable.
import socket

def connect_any(host, port, connect_timeout=3.0):
    errors = []
    for family, kind, proto, _, addr in socket.getaddrinfo(host, port, type=socket.SOCK_STREAM):
        s = socket.socket(family, kind, proto)
        s.settimeout(connect_timeout)
        try:
            s.connect(addr)
            s.settimeout(None)
            return s                      # first address whose path works
        except OSError as e:
            errors.append((addr, e))
            s.close()
    raise ConnectionError(f"all addresses failed: {errors}")

Anycast: one address, many sites

CDNs, public DNS resolvers and many API front doors use anycast: the same prefix is announced from many locations, and BGP delivers each user to whichever site their network's best path leads to. The set of users that land on a site is its catchment. Catchments follow routing policy, not geography, so a user in Lisbon may land in Frankfurt or New York, and a change at one ISP can move a whole country between sites.

Anycast gives fast regional failover: stop announcing from a site and its users move elsewhere within the propagation time. It also means that if a route change moves a client mid-connection, the new site has no TCP state and the connection resets. That is rare for short requests but matters for long-lived WebSockets and large downloads, so clients should reconnect cleanly. See anycast architecture and CDN architecture for how operators manage catchments and drains.

Hijacks, leaks and what they do to DNS and TLS

A hijack is an announcement of your prefix, or a more specific part of it, by a network that has no right to it. A leak is a legitimate route passed to neighbours who should not receive it, often a customer re-announcing its providers' routes to another provider, which can pull large volumes of traffic through a network that cannot carry it.

Two incidents show the effects. In 2008, Pakistan Telecom announced a more-specific prefix covering YouTube's addresses to block it domestically; the announcement leaked upstream and drew much of the world's YouTube traffic for about two hours. In April 2018, attackers announced more-specific routes for part of Amazon Route 53's address space, answered DNS queries for a cryptocurrency wallet site with their own server's address, and served a certificate browsers did not trust; users who clicked through the warning lost funds.

TLS with certificate validation protects you from a hijacker reading or altering traffic, provided clients never skip validation. It does not protect availability, and a hijacker who controls traffic to your DNS or web servers for long enough may be able to pass a certificate authority's domain validation. Defences are mostly in the network, but you can ask for them:

  • RPKI ROAs. Publish Route Origin Authorizations stating which AS may originate your prefixes and at what maximum length. Networks that perform route origin validation drop announcements that contradict them. If you own address space, this is the most effective single step; if you use a cloud provider's addresses, the provider does it.
  • Monitoring. Alert when a new origin AS or a more-specific prefix appears for your space.
  • CAA records and DNSSEC narrow which certificate authorities may issue for your names and make forged DNS answers detectable by validating resolvers. Read more in DNS architecture.

The 2021 Facebook outage, read correctly

On 4 October 2021, Facebook, Instagram and WhatsApp were unreachable for about six hours. Meta's post-mortem says a command meant to assess backbone capacity disconnected its data centres from each other. Its authoritative DNS servers were built to withdraw their own BGP announcements when they could not reach the data centres, a sensible health check for one failed site. Because every site failed at once, every DNS server withdrew, the names stopped resolving everywhere, and the tools engineers needed to fix it depended on the same network. The lesson is about coupling: health-based withdrawal is good, but it must not be able to withdraw everything, and you need an out-of-band path to your control plane.

BGP inside your own infrastructure

You may meet BGP closer to home. In Kubernetes, bare-metal load balancers such as MetalLB in BGP mode announce service IPs from nodes to the top-of-rack routers, and network plugins such as Calico can peer with the fabric to advertise pod and service routes. A node that dies stops announcing and the routers stop sending it traffic, so failover speed depends on BGP timers and BFD, not on Kubernetes. Cloud interconnects, such as dedicated links from a data centre into a cloud VPC, and many site-to-site VPNs, exchange routes with the cloud over BGP sessions. Private AS numbers, 64512 to 65534 and 4200000000 to 4294967294, are reserved for this kind of internal use. When an interconnect route disappears, traffic may silently fall back to a VPN path with different latency and MTU; alert on it.

Worked example: "Brazil cannot reach the API"

At 14:05 error rates for clients in Brazil jump; everywhere else is normal. Application logs show nothing because requests never arrive. Splitting connection errors by client ASN shows two Brazilian ISPs account for 90 percent. RIPEstat shows the expected origin, RPKI valid, visibility slightly down. A looking glass at one affected ISP shows its path now goes through a transit network that, according to that network's public status page, has a fibre cut in São Paulo.

There is no hijack and nothing to fix in the application. The team asks their provider to steer around the affected transit and the provider adjusts its announcements, which takes effect in about four minutes. Meanwhile, mobile clients with 3-second connect timeouts and a fallback hostname in another prefix recover on their own; partner integrations with 60-second timeouts on one address fail until the route changes.

Failure modes and trade-offs

SituationWhat users seeWhat helps
Path withdrawn at one transitErrors from some ISPs onlyMonitoring by client ASN; multi-address clients
Slow convergenceMinutes of timeouts, then recoveryShort connect timeouts, fresh connections, backoff
Hijack of your prefixTraffic to the wrong place, certificate errorsROAs, origin monitoring, strict TLS validation
Route leak upstreamHigh latency, loss, partial outageProvider escalation; little you can do locally
Anycast catchment shiftUsers served from a distant siteMeasure latency by client network; ask provider about traffic steering
Health-based withdrawal of everythingTotal outageLimits on simultaneous withdrawal; out-of-band access

The main trade-off for developers is how much routing resilience to buy. One provider's addresses are simplest; multiple providers or your own address space with multiple transits survive provider-level routing failures but bring BGP operations, ROAs and monitoring into your team's remit.

What to do next

  1. List the prefixes and origin ASes behind your public endpoints and check each with RIPEstat for visibility, origins and RPKI status.
  2. Set up an alert for a new origin AS or more-specific announcement of your prefixes.
  3. Split connection error rates by client ASN and country so partial routing failures show up as a pattern.
  4. Audit your HTTP clients: short connect timeouts, multiple addresses, fresh connections on retry, retry budgets.
  5. If you run anything that withdraws routes on failed health checks, prove it cannot withdraw from every site at once.
  6. Learn the internals next in BGP architecture, in depth and TCP behaviour under loss in TCP architecture.
Key takeaway: BGP is how the internet agrees, network by network, where your addresses are. Because each network chooses its own path, routing failures are usually partial and invisible in application logs, and recovery takes seconds to minutes. Developers cannot fix routing, but they can see it, by monitoring from many networks and checking public route data, and they can survive it, with short connect timeouts, multiple addresses, fresh connections on retry, RPKI-protected prefixes and health-based withdrawals that cannot take everything down at once.