Google Cloud at its flagship Next event introduced new cloud infrastructure enhancements that support the rising demand for artificial intelligence (AI). The infrastructure updates include the hyperscaler's latest generation of tensor processing units (TPU), the upcoming availability of its A3 supercomputer powered by Nvidia H100 GPUs, Google Kubernetes Engine (GKE) for enterprises and a new global networking platform.

“For the last two decades, Google has really been at the forefront of AI innovation – the foundation of which is, of course, our infrastructure,” Mark Lohmeyer, Google Cloud VP and GM of compute and machine learning infrastructure, told a group of reporters. AI continues to represent a major tenant of the cloud provider's workload-optimized infrastructure strategy, Lohmeyer explained.

AI workloads require “an integrated and optimized software stack” that works in conjunction with “purpose-built hardware” to support “the entirely new level of computational demands that we're seeing,” he said, noting the adoption of AI is a “once-in-a-generation inflection point in computing.”

To that point, Lohmeyer touted Google Cloud's extensive partner ecosystem – and its work with Nvidia, in particular – as enabling the hyperscaler to provide “the most comprehensive AI infrastructure portfolio of any cloud.”

Google first introduced its A3 supercomputer powered with Nvidia H100 GPUs in May and today announced A3 VMs will be generally available next month. A3 trains, tunes and serves “incredibly demanding and scalable generative AI workloads and large language models [LLMs]” at three-times the speed and 10-times the networking bandwidth of Google's previous A2 supercomputer, Lohmeyer claimed.

A3 integrates a number of Google tools including networking technologies like infrastructure processing unit offloads to support “the massive scale and performance that these workloads require,” Lohmeyer added.

In a move to offer customers “even more choice,” Google Cloud also introduced its cloud TPU v5e, which Lohmeyer claims is “the most cost-efficient and accessible cloud TPU to date.” TPUs are what Google calls its custom ASICs designed to accelerate AI and ML workloads. The company says its TPU v5e offers two-times the training performance per dollar and two-and-a-half-times the inference performance per dollar compared its previous generation.

“We're really focused on making this a scalable solution,” Lohmeyer added. “We design things across software and hardware. In this case,  the magic of that software [and] hardware working together with new software technologies like multi-slice, we're enabling our customers to easily scale their AI models beyond the physical boundaries of a single TPU pod or a single GPU cluster,” he explained. “In other words, a single, large AI workload can now span multiple physical TPU clusters, scaling to literally tens of thousands of chips – and doing so very cost effectively.”

Optimizing multicloud, multi-cluster environments for AI

GKE is another popular tool Google Cloud customers turn to for deploying and managing cloud-native AI or ML applications. The hyperscaler today introduced new capabilities for GKE enterprise edition that support multi-cluster horizontal scaling in addition to the Kubernetes orchestration platform's existing auto-scaling, workload orchestration and automatic upgrade features for general purpose compute.

The new GKE capabilities are available for cloud GPUs and TPUs, and customers are seeing significant improvements as a result. Lohmeyer cited up to 45% increases in productivity and up to 70% reductions in the time it takes to deploy software.

The hyperscaler also launched a new platform, Cross-Cloud Network, focused on simplifying secure connectivity for increasingly common hybrid or multicloud environments. The platform aims to ease operational requirements of running complex environments and help customers focus on running their business and applications rather than operating their disparate cloud network.

Lohmeyer claimed the multicloud networking platform reduces network latency by up to 35%, reduces total cost of ownership by up to 40% and offers 20-times higher threat protection efficacy compared to “other environments.”