Clockwork.io, the startup developing software to keep AI workloads running through infrastructure failures, raised $31 million in new funding.
The Palo Alto, California-based firm develops what it labels “software-driven AI fabrics,” effectively a programmable layer between hardware and workloads aiming to keep accelerators fully utilized.
Premji Invest, Wing Venture Capital, and Seligman Ventures co-led the funding round, which included participation from e& Capital and New Enterprise Associates (NEA). The raise brings Clockwork’s total funding to $73 million, with the latest capital injection going toward accelerating rollout of its suite and scaling delivery through cloud partners.
"Failures are inevitable at AI scale. Losing hours of useful work to them should not be," Clockwork.io CEO Suresh Vasudevan noted. "Fault tolerance is a goodput multiplier: it keeps GPUs doing useful work instead of waiting for recovery or repeating work already done.”
Clockwork.io was founded by Balaji Prabhakar, Deepak Merugu, and Yilong Geng in 2018, with a focus on keeping distributed computing systems running in the event of inevitable hardware catastrophes. Its suite of offerings includes LinkPass, which reroutes traffic around failed links, and TorchPass, which skirts packets away from failing GPUs.
The startup also debuted platform snapshots for TorchPass, which allow engineers to restore entire jobs from the moment prior to a failure in the event an issue is too large to migrate around.
Among those using the technology is the world’s favorite social media site LinkedIn, which has used LinkPass to skirt network fault tolerances across its AI infrastructure.
“Before Clockwork, one InfiniBand network interface card (NIC) flap could remove an eight-GPU server from service, while a switch port flap could drain a second server, doubling the impact to 16 GPUs," Raghu Hiremagalur, SVP and CTO of Infrastructure at LinkedIn, explained. “In aggregate, Clockwork.io prevents tens-of-thousands of GPU-hours of downtime per month across our fleet. By turning what were once disruptive operational incidents into manageable maintenance events, Clockwork.io has helped improve infrastructure utilization and operational efficiency.”
Distributed infrastructure operator WhiteFiber also uses the startup’s software, with plans to expand deployments across its growing global GPU-as-a-service footprint, the startup confirmed.
“Clockwork.io's automated fleet audit validates every link and node at once, localizes faults in minutes, and lets us correct them before acceptance. We bring clusters up faster, and a customer's first training run lands on a fabric validated end-to-end, not just powered on," WhiteFiber CTO Tom Sanfilippo noted. "With market demand growing as rapidly as it is, getting validated capacity to customers quickly is critical to our business, and it is why we are expanding Clockwork.io across our clusters."
Clockwork.io also claims neoclouds Together AI, Nebius, and NScale along with banking giants Wells Fargo as users.
“We built our software alongside enterprises and cloud providers operating some of the largest GPU fleets, so it handles the failures they actually see. That protection belongs in the infrastructure enterprises and cloud providers rely on every day,” Vasudevan added.
Comments