With video now ubiquitous across enterprise networks, real-time performance is more critical than ever. Yet poor network performance remains a concern due to problems such as packet loss, jitter, buffering and networks becoming far more complex.

Real-time traffic depends on factors such as the tolerance level of users for delays. Gamers and musicians dealing with real-time interaction, for example, as well as those in some areas of process control might put up with no more than 4ms of latency. In video and audio conferencing, you might get away with 20ms before you run into garbled audio or video buffering.

“Different apps perform differently and have various tolerances for latency,” said Terry Slattery, a networking consultant for NetCraftsmen, speaking at the Enterprise Connect 2023 show in Orlando, FL in late March. “You have to know what traffic is on the network and where it is traveling to, as well as the underlying application requirements.”

Here are some of the top network problems related to network performance challenges, and how to fix them. They include the following:

  1. Error rates causing packet loss
  2. Congestion
  3. Bufferbloat
  4. Latency
  5. Poor Wi-Fi design and implementation
  6. Growing network complexity
1. Error rates causing packet loss

Packet loss can happen due to various factors. This includes errors like a duplex mismatch, dirty fiber, poor fiber splits causing optical losses, pinched or long cabling, misconfigurations, and errors in the modules that plug into sw

Slattery gave the example of an optical 1Gb campus link that experienced a small number of errors per day. The error volume gradually increased over several months. Had this trend continued, it would have led to severe packet loss. Investigation revealed the cause: a bad patch cable. When fixed, the errors disappeared.

“Very small error rates have a big impact on TCP-based apps,” said Slattery. “On a 1Gb link that is traveling 1000 miles or more, you need to keep the error rate very low.”

2. Network congestion

Congestion can also cause packet loss. Network equipment generally utilizes buffers to absorb packet bursts. These buffers have a limit of the number of packets they can contain. If that limit is exceeded, lost packets result. This type of congestion is common among networks.

“Microbursts can bring about congestion on even the very best networks,” said Slattery. “Very brief periods of congestion are possible on high-speed links. Although infrequent, it does impact applications.”

One cause is a speed mismatch on links such as a 10GbE link feeding a 1Gb link. Alternatively, there can be oversubscription on links feeding a server. A good solution is to use QoS to prioritize traffic based on its characteristics. Applications generate different types of traffic and problems may be caused by a specific kind of traffic. QoS prioritization entails the queuing of packets and letting certain queues through first. It should only be applied to traffic that is causing the congestion.

Take the case of a remote site using a T3 network connection with user complaints about sluggish applications. 50% of the traffic was found to come from three websites. A small portion of it was legitimate traffic from content delivery networks (CDNs). But much of it was from entertainment sites such as music stations and Netflix downloads. With these rogue apps and sites assigned as low-priority traffic, though, problems remained. The default setting of 64 buffers per packet was increased to 256 on one or two QoS queues. Slattery said you can only do this on a few queues on the network otherwise you end up with too much buffering and overall worsening performance.

He said to review other types of congestion on networking reports. NetFlow, for example, can highlight who is using what, as well as bottlenecks at access points, on LAN to WAN routers, and between the data center and the enterprise network. He recommended active path testing available from the likes of Appneta, NeBeez, CatchPoint, and ThousandEyes.

3. Bufferbloat

Bufferbloat is due to too much buffering, which leads to jitter. This shows up as packets arriving too late for playback on voice or video and slows TCP-based apps. One real world example concerned large CAD files taking ages to be being transmitted to remote workstations. The business upgraded the link from its storage server to a L3 switch from 1Gbps to 10Gbps. However, the WAN continued to operate at 1Gbps. Remote load times increased and QoS didn’t resolve the issue. Investigation revealed the switch was being fed at 10Gbps and draining at 1Gbps. Many packets were being lost. The solution in this case was to configure a 1 Gbps path from end to end, which stopped excessive packet loss.

One way to detect bufferfloat is to conduct a long-running ping test to discover the average round trip times and in parallel begin a large file transfer. If the ping test results get longer during this test, bufferbloat is happening. One solution is to upgrade TCP software to provide active queue management.

4. High Latency

Latency rates are roughly 10ms per 1,000 miles one way. Some apps are chatty and a great many round trips take place in the background.

“TCP window scaling helps but sometimes you have to rewrite the apps to be less chatty,” said Slattery. “A simple way to reduce latency is to move the clients and data close together such as using CDNs for static content. You can also use geographical DNS to direct clients to the nearest application server.”

5. Poor Wi-Fi design and implementation

“Bad Wi-Fi is easy to do,” said Slattery. “Good Wi-Fi is harder and requires professional site surveys to guide design and implementation.”

The factors to consider include reflections from metal structures such as ducts, elevators, and building steel, as well as interference from other Wi-Fi nodes (your own or your neighbors). Access points must be placed correctly. Points to avoid: don’t strap an access point to a metal ceiling as reflective surfaces impact the signal; if an access point is ceiling mounted, don’t mount it on a wall as it won’t work as well; some ceilings are so high that the Wi-Fi signal is weak at device height; use directional antennas where needed in areas such as warehouses that mount on a wall and can shoot the signal down an aisle.

Additionally, the correct frequency bands should be used to avoid interference between access points. “5GHz and 6 GHz bands have more channels than 2.4 GHz so it is easier to ensure they don’t interfere with each other,” said Slattery.

6. Growing network complexity

Business networks were once about the LAN. But now we have the WAN, the web, Wi-Fi and the rise of edge networking, which takes advantage of Wi-Fi, satellite, and other networks. According to Chetan Sharma, an analyst with Chetan Sharma Consulting, the edge economy will be a worth $4.1 trillion by 2030. This has massive implications for network management,” he said.

In addition, network management and troubleshooting now has to deal with internet of things (IoT) sensors and devices, as well as containerized applications.

“As there is no particular device associated with a service, the network management systems are flying in a dark about service availability,” said Mike Mulica, CEO at Alef. “The service-device relationship is being disaggregated, and existing network management frameworks must change with new frameworks developed and implemented that span across different layers of the stack and provide coordination across multiple domains in multi-cloud environments.”