If you needed any proof that Mellanox would thrive under Nvidia’s rule, look no further than the company’s 51.2 Tb/s Spectrum-4 switch family, announced at GTC this week.
While the chipmaker’s new H100 GPUs may have garnered the lion’s share of CEO Jensen Huang’s attention during his nearly two-hour keynote Tuesday, a considerable amount of time was dedicated to the networking tech that makes scaling up this compute capacity possible.
Networking is essential to feeding high-performance compute (HPC) clusters with the data they need for their calculations, and increasingly the network is becoming a bottleneck, Kevin Deierling, Nvidia’s head of networking, told SDxCentral. "There's always this tension between compute and networking and boy, is it keeping us busy?”
Spectrum-4 is specifically designed for high-performance Ethernet networks where multiple 400 Gb/s links may be required per client. Nvidia’s newly announced DGX H100, for example, boasts eight 400 Gb/s network interfaces for a total of aggregate 3.2 Tb/s of bandwidth.
Based on technology acquired with the $6.9 billion purchase of Mellanox in 2019, Spectrum-4 effectively quadruples the throughput over the previous generation and offers 64, 800 Gb/s ports or up to 128, 400 Gb/s ports.
According to Sameh Boujelbene, senior research director at Dell’Oro Group, Spectrum-4 is the highest capacity Ethernet switch on the market today. “The Spectrum-4 is well ahead of the curve,” she said.
Nvidia was able to achieve this performance using 112 Gb/s serializer/deserializers (SerDes) on a single switch ASIC using a custom 4N manufacturing process developed with Taiwan Semiconductor Manufacturing Co. (TSMC). Nvidia claims to have achieved 51.2 Tb/s of Ethernet switching capacity from a single ASIC, and that it can do so in a 1U chassis.
“It’s a single chip that we can put into a 1U or 2U box. This is not a multi-ASIC design,” Deierling said, adding that there are only a handful of companies that can compete at the capacity and bandwidth offered by Spectrum-4.
"The footprint is impressive," Will Townsend, VP and principal analyst at Moor Insights and Strategy, said.
"On the surface it seems compelling given its performance, security, and lower power consumption," he wrote in response to questions. "It's also a great showcase for the BlueField DPU platform with its significant offload capability."
And to Townsend's point, the switch is part of a growing high-performance networking ecosystem designed to support the chipmaker’s compute infrastructure. Combined with Nvidia’s upcoming BlueField-3 data processing units (DPUs), ConnectX-7 smartNICs, and networking-equipped H100 CNX cards, Nvidia’s 400 Gb/s networking portfolio now spans its entire product stack.
“If you look at the analysts, they're saying, hey 400 Gb/s, people will use it, 800 Gb/s, people will use it, but there's always this caveat. There has to be an ecosystem,” Deierling said.
It should be noted that while Spectrum-4 does support 800 Gb/s ports, the infrastructure doesn’t yet exist to support those speeds beyond inter-switch communications.
Who Even Needs a 51.2 Tb/s Switch?“If you look at what Nvidia is doing on the compute side, on the storage side — but mainly on the compute side — they’re talking about 10x to 20x increase in capacity and performance,” Boujelbene said. “You don’t want your network to be the bottleneck.”
So while Spectrum-4's 51.2 Tb/s of capacity may grab headlines, the platform features a number of capabilities that make it particularly attractive for artificial intelligence (AI) workloads including nanosecond time precision.
Switch ASICs like the one found in Spectrum-4 are “definitely needed in the case of AI workloads and applications,” Boujelbene said.
But beyond the obvious use case in HPC and supercomputing clusters, specifically those being built on Ethernet fabrics, Deierling sees a number of potential applications where the Spectrum-4 may be attractive.
The hyperscalers will buy as many 400 Gb/s switches as they can get their hands on, he said. “Hyperscalers consume at such volume that I don’t think the challenge will be demand.”
And as confidential computing and zero-trust network architectures become more prevalent, he expects high-capacity switches with support for high-throughput network encryption to drive demand.
Spectrum-4 supports 12.8 Tb/s of MACsec and VXLANsec encryption, making it “the most secure end-to-end Ethernet networking platform in the world,” the company claims.
“The cloud will want to guarantee to their customers that ‘hey, we cannot look at your data,’” Deierling said. “I think once that gets adopted … it will then become pervasive. Everybody will have to do it, because ‘hey, I'm getting this in the cloud, I need to have this in the enterprise as well.’”
Boujelbene doesn’t, however, expect Spectrum-4 will see mainstream adoption anytime soon. “I don’t think Nvidia intends to make it for mainstream … it’s going to be mostly for specialized workloads, AI workloads, access workloads, and those mainly are the hyperscalers’ HPC environments."
Why Not InfiniBand?Spectrum-4 isn’t Nvidia’s only switch platform. The company actually has three distinct switching technologies including the Ethernet-based Spectrum, the InfiniBand-based Quantum-2, and the NVlink-based NVLink Switch System used in the DGX Pod.
Nvidia launched Quantum-2 last fall. The platform offers 25.6 Tb/s of throughput equivalent to 64, 400 Gb/s interfaces or 128, 200 Gb/s interfaces.
“The advantage that InfiniBand has is that you deploy and run it and all of the magic that keeps anything bad from happening is built into the network,” Deierling said. "We're doing the same with Ethernet … you just have to do so within the constraints of a multi-vendor ecosystem.”
Spectrum-4, by comparison, features more than double the transistors of Quantum-2 for twice the capacity, and a good bit more complexity, according to Deierling.
“With Ethernet, there’s a lot more complexity because you have a distributed routing function, and as much as people have tried to replace that with SDN, the vast majority of the data centers in the world today actually run using protocols that were developed in the '80s, whether that's BGP or OSPF,” he explained.
Making matters worse, AI workloads tend to require more bandwidth per client than in typical Ethernet networks. “What we've found with these AI workloads is you actually have a smaller number of much more powerful boxes,” Deierling said. "We may have one connection between the box so instead of how to make you know thousands and thousands of clients.”
And while remote direct memory access (RDMA) and RDMA-over-converged-Ethernet (ROCE) can help and is even supported by many switch vendors, “frankly nobody else really understands all the issues,” he added.
So rather than recognizing someone else’s switch for use in a HPC environment, Nvidia built its own.
“Clearly networking is becoming the bottleneck for these AI workloads, and I expect to see a lot of changes to deal with that, maybe even new topologies to accommodate those workloads,” Boujelbene
Comments