Nvidia unveiled Grove, an open source Kubernetes API designed for running AI inference workloads.
Clusters running AI inference workloads are becoming increasingly more complex. While technology like Nvidia’s own Spectrum-X and BlueField-4 keeps them interconnected and packets moving efficiently, managing the entire stack can be just as challenging.
Grove is designed to help developers run workloads more efficiently across thousands of GPUs. It's also available as a modular component within Nvidia’s Dynamo platform or directly via GitHub.
The API includes autoscaling components designed to aid scaling resources as a single resource on Kubernetes. It can be applied to individual components – like in the event of traffic spikes or all the way up to entire service replicas – to better manage overall capacity.
Grove gang
Traditional gang scheduling – or the allocation and coordination of coupled resources in more typical inference workloads – requires all components to be compiled into specific, rigid groups.
For more AI-focused workloads, however, disaggregated inference is considered by industry developers as the optimal approach. This requires separating the inference process into specific phases: prefill (for context processing) and decoding (for token generation. Both run on their own dedicated hardware resources – making the entire process infinitely more complex.
Nvidia’s Grove API handles that complex scheduling via PodCliqueScalingGroups – a tool that bundles tightly coupled Kubernetes pods with specific roles, enabling them to be scaled together and more efficiently.
A technical blog post from Nvidia engineers explained: “When scaling for additional capacity, Grove creates complete replicas … defines spread constraints that distribute these replicas across the cluster for high availability, while keeping each replica’s components network-packed for optimal performance.”
In less technical terms, Grove takes varying components of the inference stack, bundles them, and then orders them, creating an essentially more coordinated, system-level approach to AI inferencing.
“The scheduler continuously watches for … resources and applies gang scheduling logic, ensuring that all required components are scheduled together or not at all until resources are available,” the blog explains. “Placement decisions are made with GPU topology awareness and cluster locality in mind.
“The result is a coordinated deployment of multicomponent AI systems, where prefill services, decode workers, and routing components start in the correct order, are located closely for performance in the network, and recover cohesively as a group. This prevents resource fragmentation, avoids partial deployments, and enables stable, efficient operation of complex model-serving pipelines at scale.”
Comments