Distributed computing concept
– Getty Images

Distributed AI infrastructure company Zero Latency adopted Red Hat’s co-engineered AI Factory platform with Nvidia.

The firm, formerly known as Hyphastructure, is leveraging the platform as the enterprise Kubernetes foundation for its U.S.-wide network.

Zero Latency’s Zerogrid platform, which recently launched in a closed beta, acts as an orchestration for AI inference, routing workloads across edge infrastructure according to latency, locality, and capacity constraints.

Its use of Red Hat’s AI Factory platform provides it with a containerized foundation layer, allowing Zero Latency to manage graphic processing unit (GPU) resources across disparate locations within a consistent workflow.

“By using Red Hat AI Enterprise to manage distributed infrastructure, Zero Latency highlights how hybrid cloud technologies scale innovation without the burden of extensive resource investment,” Joe Fernandes, VP and GM of Red Hat’s AI Business Unit, noted. “We’re working with Zero Latency to help define the architecture for the future of low-latency AI applications."

Distributed computing, though not a new idea, is garnering increasing interest amid rising demand for computation closer to where applications are actually being deployed.

Compared to monolithic centralized infrastructure from hyperscalers and neocloud providers, the team at Zero Latency drew inspiration for its distributed platform from virtual power plants, aggregating resources into what it describes as a shared pool of inference capacity.

The firm contends the result is democratized access to Nvidia-level GPUs, with users able to power long-context or agentic AI applications while meeting demands related to latency or sovereignty.

the Zerogrid platform uses a “prefix residency index,” enabling inference cache data to reside across GPU memory, system memory, and storage tiers in a distributed cluster – an approach the company argues addresses bottlenecks brought on by large KV-cache memory states spread across multiple systems and storage tiers.

“We’ve believed for years that decentralized infrastructure beats centralized for the workloads that need it most. AI inference is the next domain it belongs in: machine-driven, constraint-bound, and poorly served by the centralized cloud,” Zero Latency CEO Michael Huerta noted. “Red Hat AI Enterprise gives us the containerization foundation to bring this architecture to enterprise customers, from the factory floor to the city street.”