Evolving technology has changed the face of network performance management. While network monitors and application performance management (APM) agents continue to gather data on network nodes across data centers and compute clouds, IT is now seeing an increasing emphasis on log management software that can help analyze such performance data.
The first star of the log revolution was Splunk. It found success as it addressed burdens of log management, adding machine learning (ML)-based analysis. Viewers suggest Cisco’s move this year to acquire Splunk – in a deal gauged at about $28 billion at the time of the announcement – represents an inflection point. The observability technology market sees new capabilities – but also new costs.
Today’s logs convey data that can be correlated with other data to tune performance from server to end user. Fodder for performance analytics includes network performance monitoring, application performance monitoring, end-user experience monitoring and varied other telemetry.
The logs also feed security information and event management (SIEM) systems that are more and more important as cybersecurity threats plague network operations.
Monitoring network complexityThe inside-cloud network metrics reported by cloud providers are one thing, but the actual experience of end users spanning the edges of the internet is another; and this has, in part, driven the observability movement, according to Bernd Harzog, CEO at APM Experts, who spoke with SDxCentral.
“The cloud providers tend to give you visibility in the network services that they provide, but not into everything else,” Harzog said. “But network performance management applies to everything.”
Network performance today is defined by the context of the applications that are implementing the business services an organization relies on, Harzog said. The context is a moving target, as software infrastructure morphs on a near-constant basis.
“People do Kubernetes, they do cloud. And guess what? Every time you make these kinds of technical changes, you have to factor that into how you monitor performance,” Harzog said. “So, the monitoring business is an ever-evolving one.”
That demanding evolution drives notable growth. Gartner expects the dual market for APM and observability products to reach an estimated $8.4 billion by 2026, with an 8.6% compounded annual growth rate (CAGR) between 2020 and 2026. The number could be much greater if a host of cloud services and network monitoring tools were added to the mix.
Many sides to dynamic log puzzleThe log management and analytics area is especially dynamic, offering new functionality tailored to providing targeted and general network and application views. Splunk’s upswing has been joined by a host of competitors.
- Observability players include Cribl, DataDog, Elastic, Grafana, Graylog, IBM Instana, New Relic, Sumo Logic, Riverbed and others.
- Original APM houses like Broadcom, BMC, Dynatrace and others have embraced the observability mantra, too.
- Sitting in a catbird’s seat of sorts are the cloud providers, led by AWS, Google and Microsoft, which manage observability data within their clouds, and sometimes beyond.
- Meanwhile, a variety of incumbent network performance vendors, including Cisco, NetScout, ManageEngine, SolarWinds and others, also play a big role in the market.
Security-related enhancements and cloud migrations lie behind much of the product activity in the area. Overall, the industry segment continues to address issues of cost, complexity and information volume that go hand in hand with log management’s wider use in the multicloud dominion.
More data metrics to understand“What we see is things getting more complex as you add cloud environments and SaaS [software-as-a-service] providers on top of that,” Robert Rea, CTO, Graylog, told SDxCentral. “You have even more data to ingest, to parse, and to really understand, and that is on an APM side, a network side and a security side.”
At Graylog, which began life as an open-source machine data analytics tool project, the emphasis today is on both SIEM and log management solutions. In July, it acquired Resurface.io’s API Security Solution to bolster its API security tooling, and more recently secured $39 million to fund growth.
Rea said network performance technology has gone beyond early use for troubleshooting purposes. But the tooling has expanded along with the complexity of environments, he said, exploiting centralized log management that allows operations teams to correlate and understand anomalies and data transfer between systems, and from cloud to cloud.
Among issues faced is the technical mismatch between established monolithic (or legacy) systems’ metrics and cloud-native microservices systems. Tracking and correlating the diverse data points is in some part behind Graylog’s move to acquire new API security tooling.
Log data correlation and the cloud“There is more network performance data than ever. It wasn’t ever easy to correlate. And now it is even more diverse,” said Richard Hartmann, community director of Grafana Labs, in an interview with SDxCentral.
This is a steady feature in technology evolution, he said, echoing the comments of APM Experts’ Harzog on the changing contexts of software infrastructure. In effect, performance and other metrics end up in “a huge mess” every time there is a substantial change in computing paradigms – from mainframes to cloud-native microservices and beyond.
Hartmann said Grafana Labs grew originally out of efforts to visualize time-series data – this led to the Grafana Cloud observability platform. It combines logs, visualization, traces, metrics and other capabilities. The company on November 14 acquired Asserts.AI, an effort started by former AppDynamics engineers to help correlate system component metrics, and to alert operations teams to performance trends as needed.
Said Harmann, “the less visibility you have, the more you have to just keep paying.” In this regard, Grafana has recently ventured deeper into the area of FinOps-style cloud-billing metrics analysis.
After the data delugeThe flood of modern machine log data can stymie organizations, according to Eileen Haggerty, area VP of product and solutions marketing at NetScout, who spoke with SDxCentral. Lately, log data may include user activity data for market planning, and make a journey to a big data lake or data warehouse.
That indicates the increased emphasis on an ecosystem approach to performance management that is seen through the lens of the transactions, she said.
“People use logs for a lot of reasons, including performance management. There are still people that use SNMP [simple network management protocol], there are still people that use NetFlow. There are people using synthetic transactions that focus on user experience. And, certainly, there's still a good share that are using the packet data itself for the source of truth throughout the network,” Haggerty said. “People are using things that are working for them.”
In fact, NetScout’s visibility platform has a special focus on deep packet inspection at scale. That is not surprising in a company that can trace its lineage back to the bellwether Network General Sniffer, a key instrument in the original rise of networked computing.
But log management is also part of the portfolio, along with key integrations. Earlier this year, NetScout announced integration of its nGeniusONE enterprise application performance platform with F5’s technology for monitoring custom applications.
“Network performance management is really at a crossroads for many organizations. That is due to the complexity, the number of vendors, the volume of traffic, and the importance of what they're migrating from one location to the other,” Haggerty said.
These tools are not “nice to haves” for organizations today, she said, rather they are “need to haves.”
Know thy application contextAs Harzog points out, the current move to apply ML analytics to network data requires gathering very precise data. It’s not getting easier. Already, users and vendors face a “boil the ocean syndrome” where intensely collecting all the available data can prove to be unaffordable for customers.
“Both users and vendors are caught in a trap. That is, in order for the AI [artificial intelligence] to work, they need high-fidelity, high-quality data,” Harzog said. This adds further complexity and cost, which grow as issues for IT shops.
Already, many users need to prioritize their observability spending in the face of what can be dramatically high service bills. Observability’s rise, at least in part, was due to cost benefits it offered versus existing APM pricing structures. But these days, cost is an often-cited factor limiting wider use of the newer tooling.
In Harzog’s estimation, the best path continues to be understanding applications in the context of the business’s objectives. That means careful prioritization.
Ask which aspects of a network you care about most, he said. If deep levels of log analytics are applied to something that is running, but its performance is not crucial to business success, advanced tooling may be a waste of money.
Comments