AI
– Getty Images

Zero Latency (formerly Hyphastructure) launched a closed beta for Zerogrid, a distributed AI inference platform designed to route workloads across edge infrastructure according to latency, data locality, and capacity constraints.

The distributed AI infrastructure company said Zerogrid is aimed at enterprise AI deployments where inference workloads increasingly need to satisfy multiple operational requirements simultaneously. These include low latency, sovereign AI and data residency rules, and fluctuating demand.

Zero Latency positions Zerogrid as an orchestration layer for distributed inference workloads, allowing them to be dispatched across geographically distributed compute resources rather than tied to a single cloud region or static on-premises deployment.

The architecture draws on concepts from virtual power plants (VPPs) in the energy sector, aggregating distributed compute resources into what the company describes as a shared pool of inference capacity.

“Zerogrid is the result: infrastructure designed for an inference world that the cloud was never built to serve,” Zero Latency co-founder Michael Huerta explained.

The company contends that inference workloads present different infrastructure challenges from AI model training, particularly as enterprise deployments move toward long-context, real-time, and agentic AI systems.

While hyperscalers and neocloud providers have focused heavily on centralized training infrastructure, Zero Latency said inference increasingly requires workloads to be routed according to factors such as latency sensitivity, geographic restrictions, and local capacity availability.

“Compute must come to the workload, not the other way around,” the company said in its launch announcement.

In a separate technical blog post, Zero Latency outlined an “admission engine” architecture designed to unify routing, cache retrieval, and recomputation decisions for distributed inference workloads.

The company said existing inference stacks often treat workload scheduling, cache management, and transport as separate processes, creating inefficiencies as AI workloads become larger and more distributed.

Instead, Zerogrid maintains what it calls a “prefix residency index,” tracking where AI inference cache data resides across graphic processing unit (GPU) memory, system memory, and storage tiers in a distributed cluster.

When a new request arrives, the system evaluates whether it is more efficient to route workloads to existing cached data, fetch cache segments across the network, or recompute missing inference state. The approach is designed to address growing infrastructure pressures created by long-context and multiturn AI workloads, which can generate large KV-cache memory states spread across multiple systems and storage tiers.

The company has not yet disclosed performance benchmarks, customer deployments, or production-scale metrics for Zerogrid. However, the launch reflects growing industry interest in distributed AI inference orchestration as enterprises explore edge AI deployments, sovereign AI architectures, and lower-latency AI services.

Nvidia and other infrastructure vendors have increasingly framed AI inference as a distributed “grid” problem, where compute resources are dynamically coordinated across geographically dispersed infrastructure rather than centralized in a handful of hyperscale regions.

Zero Latency said the closed beta is currently limited to a select group of enterprise, telecommunications, and DevOps platform users.