Nvidia’s Blackwell RTX Pro 6000 series server
– Nvidia

Google Cloud launched a series of virtual machine (VM) instances powered by Nvidia’s Blackwell RTX Pro 6000 series.

The G4 VMs are designed to power AI workloads, including fine-tuning and inference for language models. They can be partitioned with a single GPU capable of splitting into four isolated instances, each with its own high-bandwidth memory and compute cores. By offering multi-instance GPU (MIG) support, Google Cloud claims operators can run multiple smaller, distinct workloads concurrently off a single VM.

The offering can run AI models ranging from 30 billion to 100 billion parameters, taking advantage of quantization techniques to power highly intense AI workloads.

“The G4 VM offers a profound leap in performance, with up to nine-times the throughput of G2 instances, enabling a step-change in results for a wide spectrum of workloads, from multimodal AI inference, photorealistic design and visualization, and robotics simulation using applications developed on Nvidia Omniverse,” Google directors Roy Kim and Dai Vu wrote in a blog post.

The RTX Pro 6000 series is Nvidia’s mid-tier hardware offering, providing computing power capable of handling intense AI workloads without the need for costly data center-sized clusters.

The chipmaker recently added to the product line, unveiling 2U editions in August. These are RTX Pro servers that occupy two standard rack units and focus more on general-purpose computing and enterprise workloads.

Google Cloud opted for the mid-tier hardware to power its new VMs to tap the power of Nvidia’s Blackwell line without having to fork out for its higher-end chips. The hyperscaler already makes use of Nvidia’s Blackwell line with its DGX B200 and HGX B200 systems powering its Distributed Cloud Edge offering.

Each of its new G4 VMs comes in options providing one, two, four, and eight GPUs, with fractional options “coming soon.”

They’re set to complement Google Cloud’s A-series VMs, which run on Nvidia’s higher-end hardware, including GB200 Superchips for the A4X offering and the more cost-efficient G2 line.

The VMs are also fully integrated with several Google Cloud services, meaning users access them for services like Google Kubernetes Engine (GKE) and Vertex AI.

“G4 VMs provide the necessary infrastructure – up to 768 GB of GDDR7 memory, NVIDIA Tensor Cores, and fourth-generation Ray Tracing (RT) cores – to run the demanding real-time rendering and physically accurate simulations required for enterprise digital twins,” Kim and Vu wrote. “Together, they provide a scalable cloud environment to build, deploy, and interact with applications for industrial digital twins or robotics simulation.”

Google’s use of the RTX Pro 6000 comes after CoreWeave launched its own instances based on the hardware back in July.

Each of the neocloud’s instances features eight RTX Pro 6000 GPUs, alongside 128 Intel Emerald Rapids vCPUs, 1 TB System RAM, 100 Gb/s networking throughput, and 7.68 TB local NVMe storage.