Press Graphic - Vera Rubin Chip
– Nvidia

Nvidia’s upcoming Vera Rubin platform is set to power Google Cloud’s new bare metal virtual machine (VM) instances for AI inference.

The A5X instances, unveiled at the hyperscaler’s annual Next event, will be supported by Vera Rubin NVL72 rack-scale systems. They’re among the first VMs to be designed on the Vera Rubin architecture, which is due to start shipping later this year – despite rumblings of potential delays.

Notably, A5X will make use of Nvidia’s ConnectX-9 SuperNICs, which will be used to support Google’s recently unveiled Virgo scale-out networking fabric. That platform is capable of interconnecting 134,000 chips in a single east-west fabric, with the pair claiming Nvidia’s hardware running on Virgo can scale up to 960,000 Rubin GPUs in a multisite cluster.

“At Google Cloud, we believe the next decade of AI will be shaped by customers’ ability to run their most demanding workloads on a truly integrated, AI‑optimized infrastructure stack,” Mark Lohmeyer, VP and GM of AI and computing infrastructure at Google Cloud, explained.

Google Cloud already makes extensive use of Nvidia’s Blackwell GPUs to power a range of VM instances, including the A4s and A4Xs. This latest line, however, is focused around inferencing, with the pair claiming A5Xs offer up to 10-times lower inference cost per token and 10-times higher token throughput per megawatt than prior generation instances.

The A5X instances are also set to implement concepts from the open-source Falcon networking protocol being co-developed by Open Compute Project (OCP) members.

Google Cloud’s focus on inference with A5Xs adds to the debut of its 8t Tensor Processing Unit (TPU) inference-optimized versions of its custom silicon. The chips feature SparseCore dataflow processors, which offload data-dependent all-gather operations (where data is pulled from across the system only as specific calculations require it) to better prevent network bottlenecks.