All around the world Friday websites and services went all cattywampus as Cloudflare's backbone network suffered the technological equivalent of a stroke.

Any other week, this would have been a major — if short-lived — disruption to our increasingly internet-addled lives. But in a week punctuated by a massive Twitter hack and allegations that Russian agents had targeted COVID-19 vaccine research, you'd be forgiven for believing Friday's outage was the result of some nefarious group bent on depriving you of your hard-earned evening of Netflix.

The reality was anything but sinister, unless you consider pushing a bad config to be some sort of heinous act. "This was not caused by an attack or breach of any kind," wrote Cloudflare CTO John Graham-Cumming, in a blog post late Friday.

He explained the outage occurred while engineers worked on an unrelated issue with a segment of the company's global backbone between Newark, New Jersey, and Chicago. In an attempt to alleviate congestion on the network, engineers pushed a botched configuration file, which caused a cascade of outages across the company's backbone network.

Cloudflare's backbone connects many of the company's data centers over a series of private lines.

The outage lasted just under 30 minutes, but during that period the domain name service provider reported traffic across its backbone dropped by about 50% as the router was quickly overwhelmed.

Graham-Cumming wrote that the incident marked the first time Cloudflare's backbone network has gone down. However, the problem could have been much worse had Cloudflare's network been architected differently, he explained. "Because of the architecture of our backbone, this outage didn't affect the entire Cloudflare network and was localized to certain geographies," he wrote.

The affected locations included San Jose, California; Dallas; Seattle; Los Angeles; Chicago; Washington, DC; Richmond, Virginia; Newark, New Jersey; Atlanta; London; Amsterdam; Moscow and St. Petersburg, Russia, as well as São Paulo, Porto Alegre, and Curitiba, Brazil.

The outage also appears to have had ripple effects beyond these markets. During the outage, Mexico City, Denver, and Buffalo, New York were among the many cities reporting problems.

Preventing Future Outages

Cloudflare is already taking action to ensure this can't happen again. "We are sorry for this outage and have already made a global change to the backbone configuration that will prevent it from being able to occur again," Graham-Cumming wrote.

Cloudflare today introduced a maximum prefix-limit on the company's backbone BGP sessions. Had this been in place at the time of the outage, it would have shutdown Cloudflare's backbone in Atlanta. While this might sound like a bad thing, Graham-Cumming says that the outage occurred because the backbone didn't go down, and that Cloudflare's network is designed to function without a backbone if necessary.

The Cloudflare has also changed the BGP local preference for local-server routes, which the company claims will prevent a single location from attracting other location's traffic should a similar incident occur in the future.