Service providers are rapidly modernizing their networks in an effort to satiate enterprises' growing appetites for higher performance and superior user experiences. However, as service providers adopt virtualization and SDN capabilities, they are being held back by monitoring techniques that can't keep up,  a recent Analysys Mason white paper concluded.

In the paper, Analysys Mason analysts Anil Rao and William Nagy outline the pressures driving the adoption of SD-WAN and 5G technologies, the challenges service providers face delivering them, and the importance of real-time network monitoring on virtualized and cloud-native networks.

The authors report that enterprise cloud migrations and increasing reliance on software-as-a-service applications has introduced new challenges including degraded user experiences. This has driven service providers to adopt NFV, SDN, and cloud-native computing technologies.

Combined, these technologies allow for highly scalable networks where services, like SD-WAN, can be remotely provisioned on demand. However, the authors argue they also introduce complexity and require new monitoring techniques and philosophies to prevent disruptions. And this complexity is only expected to grow as technologies like edge clouds and 5G networks see wider adoption.

"In the dynamic [universal] CPE-based SD-WAN or the 5G network, [virtualized network functions] or cloud-native network functions, and service instances can be created and altered on demand, including dynamic-traffic flow changes based on SDN policies," the paper reads.

The Problem

According to Analysys Mason, the problem is rooted in antiquated monitoring methods and philosophies that leave service providers blind to network disruptions until it's too late.

"Traditional [Simple Network Management Protocol]-based, poll-based network monitoring techniques rely on non-real-time network performance data and are insufficient to assure services delivered over highly-dynamic networks," the paper reads.

Rather than using poll-based monitoring, real-time monitoring methods, like network telemetry, are gaining steam.

Cisco, Juniper Networks, and Arista Networks are just a few of the equipment vendors that have added support for network telemetry. However, while some of these vendors have added support for open standards like OpenConfig and the Yet Another Next Generation (YANG) data model, many still rely on proprietary models that make it difficult for service providers to generate a unified data set

Using network telemetry, service providers can stream and analyze network data from routers, switches, and firewalls for disruptions as they happen.

This, according to Analysys Mason, is key to mitigating microbursts. These spikes in traffic can cause packet pile ups that can degrade network performance.

"Without a monitoring solution that is well-suited for a dynamic network, the customer is left exposed to the effects of microbursts, and the service provider will have failed to assure the [quality of service] that customers expect of the network," the paper reads.

Rise of the Machines

Real-time monitoring only solves part of the problem. It's not enough for service providers to collect the data, they need to make sense of it too.

"Each of these data sources, supplemented with telemetry and active test data, provide a different level of visibility and performance metrics," the paper reads. "Service providers therefore need an extremely efficient way to correlate the diverse data sources to create a highly curated and clear data set for further processing."

"The importance of clean, accurate data cannot be overstated — the adage 'garbage in, garbage out' could not be more appropriate here," the paper reads.

Here, machine learning (ML) and artificial intelligence are already playing a role in data analysis.

Historical network data can be used to train machine learning algorithms to spot patterns and trigger responses.

According to Analysys Mason, this supervised application of ML is the most common technique used today, but it still relies on data scientists to set up and calibrate the algorithms. These supervised ML algorithms can also be extended to predictive models in order to identify network and service disruptions hours, days, or even weeks before they happen.

Meanwhile, unsupervised ML algorithms can be used to identify new patterns based on the information available to them rather than prior training.

"Applying ML/AI techniques will enable predictive operations and allow service providers to take pre-emptive action to prevent service quality degradations," the paper reads.