Distributed computing concept
– Getty Images

AI infrastructure provider WhiteFiber launched commercial availability for its cross-data-center networking solution to meet growing scale-across demand.

Dubbed WhiteFiber Continuum, the offering spans two QTS Data Center facilities some 51 miles apart. Connected via 12 dark fiber strands from Zayo, the firm claims it can provide 136 terabits per second (Tb/s) of bandwidth with just 0.9 milliseconds of round-trip latency.

DriveNets’ AI fabric provides the deep-buffer connectivity to connect both sites, while Weka was enlisted to provide the high-performance data storage and memory needed.

WhiteFiber CEO Sam Tabar described the offering as “the infrastructure answer to a problem the industry has lived with for years: geographic distance as a ceiling on what a cluster can do.”

“We are making it commercially available to enterprises that need AI compute to perform at scale, hold up under compliance requirements, and not break when a single site has a problem,” Tabar added.

WhiteFiber’s scale-across networking efforts look to extend compute interconnectivity beyond the cabinet. Distributed computing is attracting increasing interest as spiraling energy costs – or in some cases, a complete lack of power – is forcing operators to completely rethink deployments.

WhiteFiber Continuum connects multiple sites to form what it describes as a “single logical GPU supercluster.” The architecture lets users pool GPU resources across sites to while utilizing stranded telecom and metro assets for last-mile inference.

The idea of distributing potentially sensitive data might sound scary to some enterprises, but the operator stressed that its Continuum offering keeps workloads within originating jurisdictions, “even while failover and pooled AI compute happen across locations.”

“Because no single site is a single point of failure, it is designed so that training runs can continue even when one location goes down, rather than stalling until it’s restored,” the company said in a statement.

WhiteFiber has been iterating on the Continuum concept for some time under the working name Project Redwood. DriveNets was brought in back in July to help absorb AI traffic bursts before they cause congestion – a major issue for scale-across networking given the distances between nodes.

“WhiteFiber’s Continuum shows what becomes possible when ambitious engineering is paired with the right networking architecture,” Yossi Kikozashvili, DriveNets’ VP of product and GTM for AI infra, said. “DriveNets’ Ethernet-based AI fabric delivers the highest performance even in the most demanding, high-bandwidth, low-latency environments, ensuring scale-across superclusters move data efficiently, maximize GPU utilization, and optimize power efficiency.”

Interestingly, WhiteFiber said in its Continuum launch release that it has submitted patent applications covering the underlying implementation.

WhiteFiber is one of several players looking increasingly at scale across networking, with other in-action deployments already live, including Amazon Web Services’s Rainer and Microsoft’s Fairwater, though these are largely distributed workloads across disparate campuses, rather than tens of miles.

Cisco, Arista, and Nvidia are also racing to get ahead in the burgeoning market, with the latter’s networking chief recently telling SDxCentral that off-the-shelf deep-buffer switches for scale-across are “a disaster” for distributed AI workloads.