BGP and Autonomous Systems
BGP and Autonomous Systems
Definition: Border Gateway Protocol (BGP) is the path-vector routing protocol that exchanges reachability information for IP prefixes between Autonomous Systems (AS), independently administered networks, forming the routing backbone of the global internet.
How It Works
An Autonomous System is a network or group of networks under a single administrative entity with a unified routing policy, an ISP, a university, a cloud provider, a large enterprise. Each is identified by an Autonomous System Number (ASN), a 16-bit (0-65535) or, since exhaustion, 32-bit number assigned by a Regional Internet Registry (ARIN, RIPE, APNIC, etc.). Example: AS15169 is Google, AS16509 is Amazon.
BGP runs in two modes:
- eBGP (external BGP): between routers in different ASes, at the actual internet exchange points and peering links.
- iBGP (internal BGP): between routers inside the same AS, to distribute externally learned routes consistently across the AS’s own network.
A BGP speaker doesn’t just advertise “I can reach 8.8.8.0/24”, it advertises the full AS path: the ordered list of every AS a route has traversed. When AS100 hears a route from AS200 that lists path [200, 300, 400], it knows exactly which ASes the traffic will cross, and prepends its own number before re-advertising it onward.
Route selection isn’t shortest-path. BGP picks a “best path” using a decision process that weighs, in order: local preference (policy), AS path length, origin type, MED (multi-exit discriminator), eBGP over iBGP, and lowest router ID as a tiebreaker. Business relationships (peering vs transit, paid vs settlement-free) heavily shape which routes an AS chooses to prefer and announce.
Under the Hood
BGP runs over TCP port 179, a deliberate design choice: it offloads reliable, ordered delivery to TCP instead of reinventing it. Peers establish a TCP connection, then exchange messages of four types:
| Message | Purpose |
|---|---|
| OPEN | Negotiate session parameters (version, ASN, hold time, capabilities) |
| UPDATE | Advertise new routes or withdraw previously advertised ones |
| KEEPALIVE | Sent periodically to prove the peer is alive (BGP has no built-in Hello like OSPF) |
| NOTIFICATION | Report an error and close the session |
A BGP UPDATE message carries Network Layer Reachability Information (NLRI), the prefix and length, e.g. 1.1.1.0/24, plus path attributes: AS_PATH, NEXT_HOP, LOCAL_PREF, MED, ORIGIN, COMMUNITY. Communities are tags operators attach to routes to encode policy (e.g. “don’t advertise to peers”, “prefer this path regionally”) that downstream ASes can act on.
BGP is a path-vector protocol, not link-state or distance-vector: it doesn’t flood full topology (like OSPF) and doesn’t just count hops (like RIP), it propagates the actual sequence of ASes traversed, which both prevents routing loops (a router rejects any route whose AS path already contains its own ASN) and lets operators apply rich policy per neighbor.
RPKI (Resource Public Key Infrastructure) is the modern mitigation for hijacks: route origin authorizations (ROAs) cryptographically bind a prefix to the ASN authorized to originate it, letting routers reject announcements that don’t match.
Why It Matters
There is no central authority routing internet traffic. BGP is how tens of thousands of independently operated networks agree, dynamically and without central coordination, on how to reach each other. When a major AS misconfigures or withdraws routes, the effect is immediate and global, which is why BGP incidents (hijacks, leaks, outages) regularly make the news as “the internet is down” events.
BGP Path Selection Order
When multiple routes to the same prefix are received, a BGP speaker picks the best one using this priority order (simplified, first match wins):
- Highest Local Preference (local policy preference, iBGP only).
- Shortest AS_PATH length.
- Lowest origin type (IGP-learned preferred over EGP, over incomplete/redistributed).
- Lowest MED (Multi-Exit Discriminator, a hint from the neighboring AS about its preferred entry point).
- eBGP-learned routes preferred over iBGP-learned.
- Lowest IGP metric to the next hop.
- Lowest router ID, as a final tiebreaker.
This ordered list is why two networks can see genuinely different “best” paths to the same destination, each is optimizing for its own local policy first, and only falling back to path length or router ID once policy doesn’t decide it.
Common Pitfalls
- BGP hijacking: an AS announces a prefix it doesn’t own (accidentally via fat-fingered config, or maliciously), and because BGP trusts announcements by default, other networks may start routing that traffic to the wrong place. The 2008 Pakistan Telecom incident that took YouTube offline globally is the canonical example.
- Route leaks: an AS re-advertises routes it learned from one provider to another in violation of its intended role (e.g. leaking a full transit table into a peering session), causing traffic to take unintended, often congested paths.
- Assuming BGP picks the “fastest” or “shortest” path, it optimizes for policy and AS-path length, not latency or bandwidth. A shorter AS path can be slower in practice.
- Slow convergence: BGP can take tens of seconds to minutes to reconverge after a failure, since updates propagate hop by hop and route flap damping can suppress unstable routes further.
- Confusing eBGP and iBGP behavior, iBGP routers by default don’t re-advertise iBGP-learned routes to other iBGP peers (the “split horizon” rule), requiring route reflectors or full-mesh iBGP to avoid black holes.
- Forgetting that BGP sessions themselves depend on lower-layer reachability, a BGP peering session over a link that flaps intermittently produces route flapping and repeated withdrawal/re-announcement cycles that can destabilize routing well beyond the two directly affected ASes.
- Under-provisioning router memory/CPU for full routing table growth, the global BGP table has grown from a few hundred thousand entries in the early 2000s to well over a million prefixes today, and older or misconfigured hardware has historically hit hard limits (notably a widely reported 512K-route ceiling on some older routers around 2014).
Notable Real-World Incidents
- Pakistan Telecom / YouTube (2008): Pakistan Telecom tried to block YouTube domestically by announcing a bogus, more-specific route for YouTube’s prefix internally, but the announcement leaked to upstream providers and propagated globally, making YouTube unreachable worldwide for about two hours.
- China Telecom (2010): a Chinese ISP briefly announced routes for roughly 15% of the internet’s prefixes, including US government and military networks, redirecting traffic through China for about 18 minutes.
- Google/Cloudflare route leak via a small Verizon customer (2019): a misconfigured optimizer at a small ISP leaked routes into Verizon’s network, which failed to filter them, causing major outages for Cloudflare and others for roughly two hours.
- These incidents are the concrete motivation behind RPKI adoption, route filtering, and Mutually Agreed Norms for Routing Security (MANRS), an industry effort to get networks to commit to baseline route-hygiene practices.
History: ASN and Protocol Evolution
- BGP-1 through BGP-3 (1989-1991) were early, quickly-superseded iterations; BGP-4 (RFC 1771, 1994, later RFC 4271) is the version still in use today, and the one that introduced CIDR support, letting BGP carry arbitrary-length prefixes instead of legacy class-based network boundaries.
- ASNs were originally 16-bit (0-65,535), enough for the relatively small number of networks on the early internet. By the mid-2000s that space was projected to run out, so RFC 6793 (2012) extended ASNs to 32-bit, allowing over 4 billion possible numbers, with a defined transition mechanism (
AS_TRANS, using reserved ASN 23456) so old and new routers could still interoperate during the changeover. - Route filtering and RPKI adoption accelerated specifically in response to a string of public incidents (see below), moving from “best practice nobody enforces” toward increasingly common default behavior at major transit providers.
Comparison
| BGP | OSPF | RIP | |
|---|---|---|---|
| Scope | Between ASes (also within, as iBGP) | Within a single AS | Within a small AS |
| Algorithm type | Path-vector | Link-state | Distance-vector |
| Metric | Policy + AS path length | Cost (bandwidth-based) | Hop count (max 15) |
| Convergence | Slow (seconds-minutes) | Fast (sub-second to seconds) | Slow, prone to loops |
| Transport | TCP/179 | Runs directly over IP (proto 89) | UDP/520 |
Debugging a Routing/Reachability Issue
When traffic to a destination seems to take a strange path, times out partway, or a whole prefix suddenly becomes unreachable:
traceroute <destination>(ormtrfor a continuously updating view) shows the hop-by-hop path and where it stalls or loops.- A public looking glass (e.g. bgp.he.net, or a Tier-1 provider’s own looking glass) lets you query what route a different vantage point sees to the same prefix, useful for spotting whether an issue is local or global.
whois -h whois.radb.net AS<number>or a BGP looking glass shows the AS path and origin ASN currently associated with a prefix, letting you compare it against the expected owner, a mismatch is a hijack or leak signal.- Services like BGPStream or a provider’s own status page often confirm within minutes whether a widely reported outage traces back to a specific BGP announcement/withdrawal event.
- For your own announced prefixes, RPKI validators (e.g. Cloudflare’s
rpki.cloudflare.comtools) confirm whether your ROAs are correctly published and would actually be honored by validating routers.
Example
whois -h whois.radb.net AS15169 or a public looking glass shows Google’s announced prefixes. Tools like traceroute combined with a BGP looking glass (e.g. bgp.he.net) let you see the actual AS path your traffic takes across the internet, hop by hop, AS by AS.
FAQ
Does BGP know about network latency or congestion? No. Standard BGP path selection has no concept of real-time latency or link utilization, it’s driven by static policy and path length. Operators who want latency-aware routing need separate tooling (traffic engineering, anycast, CDN-level routing) layered on top.
Can a single AS have multiple ASNs? Yes, large organizations sometimes run multiple ASNs for different regions, subsidiaries, or historical/legacy reasons, all under one administrative umbrella.
What stops any network from just announcing any prefix it wants? In practice, upstream providers are supposed to filter customer announcements against an agreed prefix list (an IRR-registered route object), and RPKI lets routers cryptographically validate that an ASN is authorized to originate a given prefix. Enforcement is inconsistent across the internet, which is why hijacks still happen.
Why do some traceroutes show unexpected, longer paths? Because BGP optimizes for policy and business relationships, not shortest physical distance, a “longer” path by hop count can be cheaper or contractually preferred by one of the ASes along the way.