Researchers at Alibaba Cloud have developed a smart memory management system, Eigen+, that improves memory utilization in cloud database clusters without compromising service availability.
Instead of relying on complex time-series forecasting, Eigen+ uses machine learning to classify database instances as either transient (with unpredictable spikes) or steady (with stable memory use). Only the steady instances are considered safe for memory over-subscription, dramatically reducing the risk of Out of Memory (OOM) errors.
It applies the Pareto Principle, or the 80/20 rule, a concept that suggests 80 percent of consequences come from 20 percent of the causes.
By suggesting OOM errors adhere to the Pareto Principle, in that the majority of failures stem from a minority of instances, the researchers were able to transform a previously complex prediction problem into a classification task by simply identifying and managing problematic instances.
Put simply, instead of trying to predict when memory spikes will occur, Eigen+ identifies which database instances are prone to spikes and excludes them from memory over-subscription entirely.
The approach, first reported by The Register, echoes Quality of Service (QoS) provisioning in NFV environments where tighter resource control enables more efficient packing of services without compromising performance.
Eigen+ applies a similar logic, using classification, bin-packing algorithms, and dynamic thresholds to safely increase host utilization.
The researchers reported a 36 percent increase in memory allocation efficiency and zero OOM incidents in production deployments across MySQL clusters.
“By shifting the problem from prediction to classification, Eigen+ leverages the Pareto Principle to focus on the minority of instances responsible for the majority of OOM errors,” the researchers wrote.
Alibaba Cloud used the paper to swipe at rival hyperscalers, contending that their resource oversubscription strategies and applications, such as Google Autopilot and AWS Aurora, “often fall short in providing precise predictions, particularly in high-utilization environments where slight forecast errors can result in critical failures.”
Alibaba argues that its Eigen+ is a more robust alternative, as it sidesteps the need for fragile forecasts entirely. Instead of predicting future usage, Eigen+ focuses on identifying which workloads are safe to over-subscribe in the first place.
The paper outlining Eigen+ is a follow-up from earlier Alibaba Cloud research on Eigen, a large-scale cloud-native cluster management system for cloud databases.
The prior Eigen paper details a hierarchical resource management system capable of improved resource optimization in cloud databases, with reported improvements in allocation ratios in large-scale public-cloud production environments by over 27 percent.
Comments