Cloudflare experienced a massive outage early Sunday morning that lasted nearly five hours for some customers. In a blog post Sunday night, Cloudflare CEO Matthew Prince pinned the blame on CenturyLink, which experienced a simultaneous network-wide failure.

The cause of the CenturyLink outage remains unclear, but what is known is at 6:03 a.m. EST, Sunday Cloudflare's monitoring systems observed an increased number of 522 Errors. According to Prince, these indicated there was a problem with the connection between Cloudflare's network and where customer's applications were hosted.

These errors were isolated to CenturyLink's network, and Cloudflare's automated systems attempted to reroute traffic to alternative network providers including Cogent, NTT, GTT, Telia, and Tata. For many Cloudflare customers, the outage was resolved in less than 10 minutes as the company rebalanced traffic across its other partner networks.

"Our automated systems immediately kicked in to attempt to reroute and rebalance traffic across alternative network providers, causing the errors to drop in half immediately and then fall to approximately 25% of the peak as those paths were automatically optimized," Prince wrote.

Lingering Effects

Prince noted that due to the size of CenturyLink's network, which is among the world's largest, it was impossible to resolve the issue for all customers without intervention from network operators.

"Many hosting providers only have single-homed connectivity to the internet through their [CenturyLink's] network," he wrote. "To use the old internet as a 'superhighway' analogy, that’s like only having a single offramp to a town. If the offramp is blocked, then there’s no way to reach the town."

Making matters worse, Cloudflare reports CenturyLink's network was not honoring route withdrawals, Prince added. This meant that while Cloudflare had automatically disabled CenturyLink's network in 48 cities, the network operator was still attempting to send traffic to and from Cloudflare's.

"In the case of customers whose only connectivity to the internet is via CenturyLink/Level(3), or if CenturyLink/Level(3) continued to announce bad routes after they'd been withdrawn, there was no way for us to reach their applications and they continued to see 522 errors until CenturyLink/Level(3) resolved their issue around [10:30 a.m. EST]"

However, it should be noted that Cloudflare didn't mark the issue as resolved until 12:12 p.m. EST, Sunday.

What Happened?

It remains unclear what happened to CenturyLink's network to cause the outage. In an email to SDxCentral, CenturyLink provided little detail in regard to the cause of the outage. A spokesperson for the network operator simply stated that "on Aug. 30, customers in several global markets were impacted by an IP outage across the network. All services have been restored."

SDxCentral has requested additional details into the cause of the outage. This story will be updated as more information becomes available.

In Cloudflare's blog post, Prince speculated on the cause of the outage. He wrote that Cloudflare observed a large spike in border gateway protocol (BGP) updates on CenturyLink's network around the time of the initial outage.

This was backed up by an update provided by CenturyLink that stated a bad flowspec rule — which is used to distribute firewall rules using BGP — may have caused the issue.

"Flowspec is a powerful tool," wrote Prince. "It is great when you are trying to quickly respond to something like an attack, but it can be dangerous if you make a mistake."

Prince added that after Cloudflare experienced a flowspec outage about seven years ago, the company stopped using the command due to the potential for large network outages.

"We can only speculate what happened at CenturyLink/Level(3), but one plausible scenario is that they issued a flowspec command to try to blog an attack or other abuse directed at their network," he wrote.

Prince speculated further that a bad flowspec command late in the BGP queue could have caused a loop that exhausted the router's memory and overwhelmed the router's CPU. This, he added, would have made it difficult for CenturyLink to access the routers and correct the issue.

However, Prince adds that there is a distinct possibility that the outage may not have been caused by CenturyLink and instead caused by one of its customers.

He wrote that if a downstream customer had used a botched flowspec rule to block an attack, it could have caused the cascade that took down CenturyLink's network.

Second Outage

The outage marks the second major outage of Cloudflare's network in a little over a month.

In July, a bad router config intended to relieve congestion on the network caused a cascade of outages across the company's backbone network as all traffic was sent to a single router in Atlanta.

The outage lasted just under 30 minutes, but during that period the domain name service provider reported traffic across its backbone dropped by about 50% as the router was quickly overwhelmed.