A team of researchers from Shanghai Jiao Tong University and Huawei has proposed a new way to share GPUs more efficiently across jobs in campus data centers, reducing idle GPU time and job wait times.
In their paper, the engineers argue that current GPU pooling solutions fall short when multiple applications compete for GPU resources – a common scenario in research environments.
It’s a regular problem in academic settings, where resources and access to high-end GPUs are limited. One campus data center analyzed by the researchers saw GPU utilization run below 25 percent while more than 200 jobs waited in the queue.
Their method, called gPooling, works by intercepting GPU driver APIs at the kernel level to slice GPU access by time and memory, allowing multiple workloads to run on a single physical GPU.
The system integrates with existing workload managers like Slurm, without requiring any changes to source code or job scripts.
In real-world tests using GPU nodes in a campus data center, gPooling improved GPU utilization by up to 2x and reduced job wait times by 21-72 percent, with the biggest gains in heavily loaded scenarios.
The researchers tested gPooling on a substantial computing cluster with 936 nodes and Nvidia HGX A100 systems, using real workloads including deep learning training, scientific computing, and large language model tasks.
The research focused on university environments, but the approach could potentially help other organizations facing similar GPU underutilization issues.
The researchers contended that existing technical approaches to GPU sharing, like Multi-Process Service (MPS), Multi-Instance GPU (MIG), and virtual GPU (vGPU), have “reliability issues and limited usage scenarios," while hardware solutions "can only provide limited elasticity.”
“By intercepting GPU driver APIs, gPooling achieves fine-grained arithmetic, storage control and state management of GPUs,” the paper reads. “gPooling achieves accurate arithmetic slicing through time-slice control and accurate memory control through interception of memory APIs… As such, gPooling provides both flexibility and generality at a low-performance overhead.”
The use of A100 GPUs in the tests is notable given that US export restrictions on high-end semiconductors meant the Chinese researchers weren’t able to test their gPooling method on more powerful hardware.
Comments