As organizations increasingly race to build and train artificial intelligence (AI) models, it places a corresponding strain on existing networking infrastructure.
AI model development and training relies on GPUs and AI-accelerator technologies that have different networking characteristics than CPUs. While closely clustered systems using InfiniBand is one option, it's not the only one.
This week, DriveNets, announced its entry into the AI networking space with its Network Cloud-AI product. It's an approach that relies on Ethernet, rather than InfiniBand, to help connect large numbers of AI-optimized systems in a distributed cluster. DriveNets claims that its new Network Cloud-AI technology approach has the capability to connect up to 32,000 GPUs together in an AI cluster, with bandwidth connectivity of up to 800G.
[ Follow SDxCentral's complete AI coverage ]"DriveNets approach is unique as it combines whitebox switching and its own software," Allan Weckel, analyst at 650 Group, told SDxCentral. "This will allow them to address the full list of cloud customers, including those who do white boxes."
The market for AI networking is potentially massive, with 650 Group sizing the market as a $10 billion by 2027. Multiple vendors are already active in the space including Arista, which is also taking an Ethernet-based approach to help enable high-speed AI optimized networking.
Taking a distributed disaggregated chassis approach to AI networkingDriveNets was founded in 2015 and has raised $587 million in funding to date. Among the company's biggest users is AT&T, which is now using DriveNets Network Cloud technology to help provide a software-based core routing technology.
The basic premise behind DriveNet's Network Cloud is that in order to scale a cloud network, an operator doesn't want to be limited to a single routing chassis.
"The whole software is cloud native, and you can run multiple different networks over the same shared infrastructure of physical white boxes," Inbar Lasser-Raab, chief marketing and product officer at DriveNets, told SDxCentral.
DriveNets has developed software that will run on whitebox hardware. The architecture is an approach known as a distributed disaggregated chassis (DDC). Lasser-Raab said that the idea is that a DDC deployment is just like a chassis, but instead of being kind of limited to one rack, it can be distributed over up to 200 white boxes.
"So it's literally the largest router in the world as a single entity," she said.
The network cloud approach was originally developed to support common data traffic, but with the new Network Cloud-AI approach it has been optimized to support large-scale AI workloads as well. Lasser-Raab said that every hyperscale provider today is building its own AI accelerator platforms. In her view, those organizations don't want to be limited to InfiniBand technology that comes from Nvidia and only supports Nvidia GPUs. She noted that network providers need a fabric to literally create a very large switch that connects to all those GPU servers. That fabric is what DriveNets is aiming to deliver with its Network Cloud-AI.
"The most expensive thing about AI infrastructure are the GPU and AI accelerators in the network," Lasser-Raab said. "Now what happens is if you don't utilize those GPUs to the maximum, you're just leaving money on the table.”
Comments