Google's headquarters in Silicon Valley in Mountain View, California - SDx Cropped
– JHVEPhoto/Getty Images

Google Cloud has added agentic AI capabilities for Kubernetes workloads with updates to the Google Kubernetes Engine (GKE).

As announced at KubeCon 2025, the updates come in response to the so-called agentic wave, which is enabling a shift from predictable, rule-based software to AI agents autonomously working on code, decisions, and tools on behalf of enterprises.

Marking 10 years in production, GKE has been updated to support 130,000-node clusters in a bid to meet the massive scale needed for foundational model training workloads.

Meanwhile, new tool GKE Agent Sandbox offers integrated sandbox snapshots and container-optimized compute. Google claimed that Agent Sandbox can fully restore the memory state of sandboxes in less than three seconds when creating snapshots of persistent volumes that GKE pods use to store data. This enables warm rather than cold starts, in which the deployment of an AI model is delayed after a period of inactivity.

While Agent Sandbox relies on Google’s own GVisor container software to isolate agentic environments, a Google Cloud spokesperson stressed to SDxCentral the tool is compliant with the Cloud Native Computing Foundation (CNCF) framework via its pluggable open-source API.

“Our strategy is to elevate the open-source core first. Agent Sandbox is being built as a CNCF project and is pluggable by design, supporting various isolation technologies beyond just our defaults,” said Google Cloud. “In addition, our participation in the new Kubernetes AI Conformance program guarantees interoperability, ensuring that AI on Kubernetes remains portable and free from vendor lock-in.”

Agent Sandbox was also said to enhance safety by providing kernel-level isolation built foundationally on GVisor, reducing the risks of data loss, exfiltration, or damage to production systems.

Google Cloud also touted the agent's simple usability for AI engineers via an API and Python software kit that “abstracts away complex Kubernetes YAML configurations into a simple context manager, allowing sandbox lifecycle management without deep infrastructure expertise.”

According to Google Cloud, the rise in agentic AI means existing infrastructure must now support orchestrating thousands of ephemeral sandboxes, rapidly creating and deleting them as needed, while strictly limiting their network access. Infrastructure also has the challenge of scaling and meeting governance requirements while managing the countless sandboxes that are scheduled in parallel.

By scaling GKE to 130,000 nodes in a single cluster, a spokesperson told SDxCentral the engine could manage the computational power needed to train massive AI models, which often require hundreds of thousands of accelerator chips.

“It also signals a move toward massive, single-cluster scalability for tightly coupled jobs. The goal is to enable platforms that can manage millions of chips while maintaining the workload fungibility and ease of management found in a single cluster,” said Google.

In addition to its tenth anniversary update of its container management service, the hyperscaler has also brought GKE Inference Gateway to general release, bringing optimizations that reportedly achieve up to 96% lower time-to-first-token (TTFT) latency.

GKE recently came out on top of Gartner’s most recent Magic Quadrant ranking for container management, with the analyst firm praising Google for “keeping its user experience simple while adding advanced features.”