BlogsCloudflareBGP Traffic Engineering

BGP Traffic Engineering

BGP Traffic Engineering

23
posts
2012–2023

Cloudflare's BGP traffic engineering capabilities have been demonstrated to automatically reroute traffic around undersea cable failures, ensuring service continuity. The platform leverages BGP to seamlessly switch to backup paths, such as from Paris to Jersey when the direct London to Jersey path is disrupted, highlighting the resilience and adaptability of the global network. This post details an incident where packet loss on a major transit provider (Telia Carrier) necessitated manual intervention, and analyzes a significant outage experienced by Virgin Media (AS5089) due to BGP withdrawals and announcements, impacting DNS resolution and overall connectivity.

2023

Cloudflare’s view of the Virgin Media outage in the UK

4/4/2023

This post analyzes the Virgin Media outage in the UK on April 4, 2023, by examining Cloudflare Radar data for traffic drops and correlating them with BGP announcement and withdrawal activity from Virgin Media's autonomous system (AS5089). It details the impact on DNS resolution for virginmedia.com due to authoritative nameservers being within the affected network and highlights the BGP activity that preceded and accompanied the traffic disruptions.

2022

Helping build a safer Internet by measuring BGP RPKI Route Origin Validation

12/16/2022

Introduced a new method to measure the deployment of RPKI Route Origin Validation (ROV) by comparing the reachability of RPKI-valid and RPKI-invalid prefixes from measurement points within an AS. This method leverages the `isbgpsafeyet.com` website and continuous measurements from end-user browsers to gather data on ROV deployment. Presented findings on the global state of ROV deployment, indicating that while RPKI signing is progressing, ROV adoption is still lagging significantly, with substantial variations across different countries.

Why BGP communities are better than AS-path prepends

11/24/2022

This post introduces BGP communities as a superior alternative to AS-path prepending for traffic engineering. It details the BGP best path selection algorithm, highlighting local preference and AS-path length. The post explains how BGP communities can be used to influence inbound traffic by providing more granular metadata than AS-path prepending, which is often overridden by other attributes.

Cloudflare’s view of the Rogers Communications outage in Canada

7/8/2022

This post details the analysis of a major outage at Rogers Communications, examining the BGP update patterns and prefix withdrawals that led to the widespread disruption, and observing the recovery process. It provides specific data points on BGP update spikes, prefix withdrawals and advertisements, and traffic levels from Cloudflare Radar, illustrating the technical impact of the outage.

Cloudflare outage on June 21, 2022

6/21/2022

This post details a specific incident where a re-ordering of terms in BGP policy statements caused the withdrawal of critical site-local prefixes, leading to a widespread outage. It explains the technical cause of the error, including the interaction between the `REJECT-THE-REST` term and the `ADV-SITE-LOCALS` terms, and the subsequent impact on internal load balancing (Multimog). The post also outlines immediate remediation steps focusing on process improvements (smaller stagger steps for MCPs), architectural redesign of policy statements, and automation enhancements (commit-confirm rollback).

Tracking shifts in Internet connectivity in Kherson, Ukraine

5/4/2022

This post details the observation of a shift in BGP routing for AS47598 (Khersontelecom) from Ukrainian networks to Russian networks (AS201776 - Miranda, AS12389 - Rostelecom) during an internet outage in Kherson, Ukraine. It also documents the subsequent return to Ukrainian networks. The post analyzes the impact of this routing change on Cloudflare's Anycast data center selection, showing a temporary shift to the Moscow data center and then back to Kyiv and Frankfurt. The analysis uses BGP routing data and Cloudflare Radar traffic data to illustrate these shifts.

BGP security and confirmation biases

2/23/2022

This post details a specific incident where a configuration error during a hot-cut maintenance operation caused a route leak of up to 2,000 Internet prefixes. The error involved the removal of BGP filters without adding explicit policies to only export Cloudflare and customer prefixes, leading to the ISP's prefix-limit protection being triggered. The post discusses the timeline of the event, the root cause analysis, and the immediate mitigation steps taken, including implementing an implicit reject policy for BGP sessions. It also highlights the importance of RFC8212 and the need for technical solutions like RPKI to prevent route leaks.

2021

Protecting Cloudflare Customers from BGP Insecurity with Route Leak Detection

3/25/2021

Introduced Route Leak Detection, a new network alerting feature that notifies customers when their prefixes onboarded to Cloudflare are being advertised by unauthorized parties. This feature monitors global routing tables for unexpected changes related to customer prefixes by ingesting data from RIPE RIS, RouteViews, and Caida's public BMP feed. It correlates these changes with Cloudflare's own actions to identify potential hijacks and aims to detect leaks within five minutes. The feature supports configuration via the Notifications tab and integrates with email and PagerDuty. It also highlights the importance of RPKI for preventing route leaks and notes that over 50% of top Internet providers now support RPKI.

2020

August 30th 2020: Analysis of CenturyLink/Level(3) outage

8/30/2020

This post analyzes a major CenturyLink/Level(3) outage, detailing Cloudflare's automated traffic rerouting mechanisms using alternative network providers to mitigate the impact. It investigates the likely root cause as a Flowspec rule that disrupted BGP announcements, leading to widespread connectivity issues. The analysis includes observations on BGP update volume spikes, the potential for a looping Flowspec rule, and the challenges in resolving the outage due to router instability and the inability to honor route withdrawals.

2019

Dogfooding Magic Transit by DDoSing our Austin office

10/24/2019

This post details the dogfooding of Magic Transit by simulating a DDoS attack on the Cloudflare Austin office. It explains how Magic Transit uses BGP to attract traffic to Cloudflare's network edge, where it is protected by multi-layered defenses including XDP, eBPF, iptables, Gatebot, and DoSD. The experiment involved onboarding the Austin office to Magic Transit in an always-on routing topology and then launching a simulated DDoS attack. The post highlights the effectiveness of the mitigation systems in blocking attacks quickly and without service degradation, and discusses lessons learned for engineering and procedural aspects, including fine-tuning mitigations, optimizing routings, and refining run-books for faster customer onboarding.

The deep-dive into how Verizon and a BGP Optimizer Knocked Large Parts of the Internet Offline Monday

6/27/2019

This post provides a detailed technical analysis of a BGP route leak event that affected large parts of the internet. It uses RIPE NCC archived data and shell commands to extract and analyze BGP updates, focusing on Cloudflare's routes. The analysis identifies the affected routes, their timing, and the role of BGP communities (specifically NO_EXPORT) in potentially preventing such leaks. It also touches upon the behavior of BGP route optimization systems and the implications of invalid route propagation.

How Verizon and a BGP Optimizer Knocked Large Parts of the Internet Offline Today

6/24/2019

This post details a major BGP route leak incident caused by a BGP optimizer product from Noction, which split IP prefixes into more-specific routes that were then improperly announced by DQE Communications and Verizon. The incident resulted in a significant loss of global traffic for Cloudflare and other services. The post explains the underlying BGP mechanism, the role of more-specific routes, and the failure of Verizon to implement best practices like prefix limits, IRR filtering, and RPKI validation. It also describes Cloudflare's efforts to resolve the incident by contacting the involved networks and advocates for the widespread adoption of RPKI.

2018

How a Nigerian ISP Accidentally Knocked Google Offline

11/15/2018

This post details a specific incident where a Nigerian ISP, MainOne, accidentally misconfigured their network, causing a route leak that led to a 74-minute outage for Google and other services. It explains the concept of route leaks, the role of Autonomous System (AS) numbers and AS Paths in Internet routing, and how the lack of proper filtering by intermediary networks (CN2, TransTelecom) allowed the misconfiguration to propagate widely. The post highlights Cloudflare's auto-remediation systems detecting and mitigating the incident, and reiterates the company's commitment to improving Internet routing security through RPKI validation and BCP-38.

BGP leaks and cryptocurrencies

4/25/2018

This post details a specific incident where a BGP leak by eNet Inc. (AS10297) announced Amazon's IP space (AS16509), affecting Route53 DNS servers and leading to potential cryptocurrency theft. It explains the mechanics of BGP leaks, the actors involved, and the impact on Cloudflare's DNS resolver (1.1.1.1) in specific regions. It also discusses potential solutions such as RPKI, DNSSEC, HSTS, and DANE to mitigate such attacks.

2016

Not one, not two, but three undersea cables cut in Jersey

11/30/2016

This post details how Cloudflare's BGP routing protocol automatically rerouted traffic from London to Paris to Jersey after three undersea cables serving Jersey were cut. This demonstrates the resilience of the network and the effectiveness of BGP in handling unexpected infrastructure failures by seamlessly switching to backup paths.

The Internet is Hostile: Building a More Resilient Network

11/8/2016

This post details the development of an automated system for monitoring internet quality, detecting brownouts, and automatically mitigating issues by disabling problematic links. It describes the use of thousands of probes to monitor packet loss, the implementation of a detection and alerting mechanism based on predefined triggers, and the development of automated mitigation actions using NAPALM and Salt. The post also highlights the impact of these improvements on reducing customer-facing errors.

A Post Mortem on this Morning's Incident

6/21/2016

This post details an incident of significant packet loss on the Telia Carrier network, impacting Cloudflare's connectivity and leading to a spike in 522 errors for customers. It describes the manual intervention of taking down ports with the problematic provider and the subsequent rerouting of traffic to other providers. The post also highlights the limitations of BGP in detecting packet loss and the need for augmented systems for proactive detection and automated mitigation. It details the plan to extend automated detection and mitigation to all POPs and improve communication during incidents.

2014

Route leak incident on October 2, 2014

10/2/2014

This post details a specific incident of a BGP route leak caused by Internexa, an ISP in Latin America. The leak directed traffic destined for Cloudflare data centers globally to a single data center in Medellín, Colombia, overwhelming its capacity and causing a 49-minute outage. The incident also disrupted traffic for Telecom Argentina due to a separate leak. The post quantifies the impact on traffic in North America (50% drop) and Europe (12% drop). It also references past route leak incidents and outlines Cloudflare's immediate response: working with Internexa to resolve the leak, quarantining the Medellín data center, and disabling connectivity with Internexa while investigating the cause. An internal post-mortem was initiated, and service credits were planned for affected customers.

2013

Today's Network Issue

5/31/2013

This post details an incident where a network migration by a transit provider (GTT/nLayer) caused a large amount of traffic to be misrouted to Cloudflare's Los Angeles data center, leading to overload and service degradation for approximately 10% of connections. The issue was resolved by working with the transit provider to rebalance traffic. Cloudflare has since added this scenario to its monitoring and mitigation strategies to prevent future occurrences.

Syrian Internet Restored

5/8/2013

This post analyzes the Syrian internet outage and restoration, providing evidence against the government's claim of a cable cut. It details the systematic withdrawal and restoration of BGP routes across Syrian providers, contrasting it with the expected behavior during a physical cable cut. It also identifies a small subset of Syrian IP space that maintained connectivity, further discrediting the official explanation. The post includes a graph of Syrian traffic to Cloudflare's network during the event.

How Syria Turned Off the Internet (Again)

5/7/2013

This post details an incident where Syria's internet access was withdrawn by systematically withdrawing BGP routes from its border routers. Cloudflare observed this event through its network data and monitoring of network routes, noting that this was the same technique used in a previous outage. The post includes a video generated by BGPlay showing the route withdrawals and a graph of Cloudflare's network traffic from Syria.

2012

Why Google Went Offline Today and a Bit about How the Internet Works

11/6/2012

This post details an incident where a route leak from Moratel caused a significant outage for Google's services. The author, a network engineer at Cloudflare, troubleshooted the issue by analyzing DNS resolution failures and BGP routes. The analysis revealed that Moratel was incorrectly announcing routes for Google's IP addresses, causing traffic to be misrouted via Indonesia. The fix involved contacting Moratel to stop the incorrect announcements. The post also explains the BGP trust model and the concept of route leakage, referencing a similar incident involving YouTube.

Today's Outage Post Mortem

5/2/2012

This post details a significant outage caused by a misconfiguration in the Hong Kong data center's BGP routing announcements, which incorrectly directed global traffic to an offline facility. It explains the mechanics of BGP routing and how the error propagated. Future prevention measures include implementing a verification layer for routing changes and enhancing confirmation checks with upstream providers.