Cisco introduced Nexus HyperFabric AI cluster solution, co-developed with Nvidia, aiming to simplify the deployment and management of artificial intelligence (AI) infrastructure for enterprises using Ethernet networks, at this week’s Cisco Live 2024.
The cloud-managed data center network solution combines Cisco's AI-native networking and Ethernet switching expertise with Nvidia accelerated computing and AI software and VAST data storage capabilities. It aims to offer a single place to design, order, deploy, validate, monitor, assure and optimize all the AI pods and data center workloads.
Kevin Wollenweber, SVP and GM of Cisco Networking, told SDxCentral, the generative AI infrastructure deployment is still in its early stages, despite the rapidly growing market. Hyperscalers like Meta are using a mix of InfiniBand and Ethernet technologies for their AI clusters, but enterprises and service providers, that typically use pre-trained models or fine-tune existing models with a focus on inference, don't have the same scale.
In addition, many enterprises are comfortable with Ethernet infrastructure and the goal of this Nexus HyperFabric solution is to help these enterprises participate in the generative AI adoption using their familiar Ethernet systems.
Cisco and Nvidia on Ethernet networking for AI Ethernet has long been the backbone of enterprise networking due to its widespread adoption and cost-effectiveness. Despite the rise of Infiniband in the generative AI era as a high-throughput low-latency fabric, Ethernet remains the default networking choice for most data centers.
Nvidia CEO Jensen Huang underscored this point, “Not every data center can handle Infiniband, because they’ve already invested their ecosystem in Ethernet for too long.” The company is pushing its Spectrum-X Ethernet platform to serve this need.
Similarly, as a major networking vendor, Cisco introduced the Nexus HyperFabric solution to integrate seamlessly with customers’ existing Ethernet networks. “These enterprises are coming to us and saying, ‘Look, we have ethernet infrastructure. We're used to Ethernet. We're used to the modern operating model of Ethernet,’” Murali Gandluru, VP of product management at Cisco, told SDxCentral.
But, the Cisco and Nvidia partnership “is much more than just retellings of Nvidia GPUs and NICs and connecting it to an existing Ethernet fabric that we've already built,” Wollenweber said. “This is really [a] new operating model, which is how do we simplify the deployment of these technologies and let the data scientists focus on running AI workloads and efficient levels across the fabrics.”
What is the Cisco Nexus HyperFabric AI cluster? Similar to the Cisco Meraki product line offering cloud management for campus and branch networks, the networking giant claims the Nexus HyperFabric AI cluster solution is designed to bring these cloud-managed operations and management to data centers.
The solution includes the following elements: The vendor is expected to roll out the solution toward the end of the second half of this year, with large-scale deployment targeted for early 2025, according to Wollenweber.
- Cisco cloud management capabilities to simplify IT operations across all phases of the workflow.
- Cisco 6000 series switches for spine and leaf that deliver 400G and 800G Ethernet fabric performance.
- Cisco Optics family of QSFP-DD modules to offer customer choice and deliver high densities.
- Nvidia AI Enterprise software to streamline the development and deployment of production-grade generative AI workloads.
- Nvidia NIM inference microservices accelerate the deployment of foundation models while ensuring data security.
- Nvidia Tensor Core GPUs starting with the Nvidia H200 NVL, are designed to supercharge generative AI workloads.
- Nvidia BlueField-3 data processing unit (DPU) DPU processor and BlueField-3 SuperNIC for accelerating AI compute networking, data access, and security workloads.
- Enterprise reference design for AI built on Nvidia MGX, a modular and flexible server architecture.
- The VAST Data Platform offers unified storage, database and a data-driven function engine built for AI.
Introducing a new operating model In addition to the launch of Cisco Nexus HyperFabric, the vendor also enhanced its Nexus dashboard to centralize operations and automation, and consolidate management of multiple fabric technologies.
“One of the key things we want to do as enterprise networking vendors is to provide customers the flexibility of choosing their operating model,” Gandluru said.
“If you have a private cloud-based operating model where you want to run your infrastructure, completely on prem, you have the Nexus dashboard,” he added. “Cisco Nexus HyperFabric is about providing an operating model that is Cisco-managed as-a-service, cloud-based management plane and the policy plane.”
Comments