Generic data center administrator image
– Getty Images

Cluster Director, Google Cloud’s infrastructure for managing high-performance systems, is now generally available as the hyperscaler looks for ways to help customers better configure their AI infrastructure.

Debuted in April, Cluster Director acts as a unified management plane aimed at making the management of large scale clusters via Slurm and Kubernetes that little bit easier.

The offering automates setup of clusters, and intuitively integrates Google Cloud solutions covering compute, networking, and storage via a single environment, with the hyperscaler claiming users can spin up standardized, validated clusters “in minutes.”

Users can leverage Cluster Director to automate lower-level operations through its control plane, API, or CLI for jobs using Slurm as well as custom orchestrators. Before workloads even touch a GPU, Google claims that Cluster Director runs a suite of health checks and performance validations to verify network and accelerator integrity.

“It replaces fragile DIY tooling with a robust topology-aware control plane that handles the entire lifecycle of Slurm clusters, from the first deployment to the thousandth training run,” a Google blog post explained.

The offering supports a range of systems, including Google Cloud’s A4X and A4X Max VMs, powered by Nvidia's Blackwell GPUs.

“[Google Cloud's Cluster Director] complements the power and performance of Nvidia’s accelerated computing platform,” said Dave Salvator, director of accelerated computing products at Nvidia. “Together, we're providing customers with a simplified, powerful, and scalable solution to tackle the next generation of computing challenges.”

Cluster Director is now generally available, while Cluster Director support for Slurm on Google Kubernetes Engine (GKE) is now in preview.

Support for Slurm comes after its lead developer SchedMD was acquired by Nvidia earlier this week.

“By running a native Slurm cluster directly on top of GKE, we are amplifying the strengths of both [researchers and platform teams],” Google’s blog post reads. “Researchers get the uncompromised Slurm interface and batch capabilities, such as sbatch and squeue, that have defined HPC for decades, [while] platform teams gain the operational velocity that GKE, with its auto-scaling, self-healing, and bin-packing, brings to the table.”