LONDON – Nutanix is pushing enterprise infrastructure control to a new level with Agent Gateway, an AI control plane that centralizes governance and optimization of model usage, while giving CIOs a way to rein in “free‑for‑all” token spending.
Positioned between users, applications, and a growing mix of open‑weight and frontier models, the gateway lets enterprises set policies on who can use what, for which workloads, and at what cost. Its introduction aims to curb unnecessary token spending, providing a centralized view for businesses of who is consuming what and how to control token usage.
Where employees today may use frontier-level models for trivial tasks like simple document summarizations, Nutanix CEO Rajiv Ramaswami touts its control plane offering as a means to define ROI from AI deployments, defining who can use which tools and models, for which use cases, and how much they are permitted to spend in tokens.
“It's a free-for-all. Anybody can go to anything they want,” Ramaswami said in a press briefing. With the vendor’s Agent Gateway, however, he explained that businesses can put in place rules that give engineering teams access to “a simple model” for a set of use cases, while reserving the fanciest system for the most difficult, multi-agent applications.
Ramaswami said the AI gateway concept is already resonating at the C‑suite level prior to launch, noting that in a dozen CIO and COO meetings during a recent trip to London, such an idea was floated and that Nutanix is now pushing partners to catch up so they can take the message to customers.
“Every meeting, this concept of gateway came up, and they were very intrigued at the CIO level, even the chief operating officer (COO) level,” Ramaswami said. “So now it's starting to be a C-suite thing, whether it's a COO or a CFO, not just a CIO thing, in terms of trying to figure out how to enable AI for the enterprise in a cost-effective way.”
Agent Gateway is part of the Nutanix AI stack (Enterprise AI 2.7) and can connect AI users and agents to models and model context protocol (MCP)-compatible tools and servers, where it enforces pre-set policies and rules defined by the infrastructure operators themselves.
Ramaswami promised the platform will be iterated upon over time, with the CEO envisioning it as an “AI within an AI.”
“Imagine if it starts getting smarter and to understand the application itself … it can then say, 'OK, well, for this particular set of use cases, I know which model it needs to go to.' I don't have to specify it. It figures out what model to go to and provides access to the appropriate model and optimizes the cost.”
For now, the gateway fronts Nutanix’s graphic processing unit (GPU)‑backed inference stack, which runs on Kubernetes and exposes shared inference endpoints for a mix of open‑weight and frontier models. While that stack is currently Nvidia-based, the vendor plans to support AMD “by the end of the year,” a moave that follows the chip giant's $150 million investment in Nutanix back in February.
Referencing that investment, Ramaswami said Nutanix ultimately wants to be hardware agnostic, extending support for multiple hardware platforms while providing inferencing stacks, and now a gateway to “enable enterprises to cost-effectively and easily deploy and consume it.”
“There's going to be a variety of choices on the underlying basis of how you optimize [clusters], and maybe you say, well, it's cheaper for me to do this on a Google TPU [Tensor Processing Unit] or on an AMD GPU compared to an Nvidia GPU. There's going to be a whole range of options,” Ramaswami added.
Comments