Amazon Web Services (AWS) has created an unconventional workaround to keep using its custom networking hardware in Nvidia's latest rack-scale systems: a separate cabinet just for network cards.
According to analyst firm SemiAnalysis, Nvidia’s flagship NVL72 rack-scale system uses shorter 1U server trays – meaning hyperscalers like AWS can’t physically fit its network cards into these slimmer trays.
For the previous generation GB200s, AWS employed a custom design called NVL36x2 to connect two 36-GPU clusters to create a 72-GPU system. That workaround, however, required extra server trays to fit some nine network cards per tray, which resulted in more bugs than in using Nvidia’s native NVL72 design.
In a bid to avoid similar issues with Nvidia’s latest rack-scale system, and to also get around Nvidia’s decision only to offer their native design for the GB300, AWS employed a new solution, dubbed JBOK – ‘just a bunch of K2v6 network interface cards,’ or NICs – a punny nod to JBOD (just a bunch of disks).
Instead of squeezing its custom NICs into smaller trays, AWS got creative. Its engineers made it so that network cables from the GPUs connect to a separate side cabinet full of just network cards.
The additional network-focused cabinet contains some 18 taller (2U) trays with enough room to hold all of the hyperscaler’s custom network cards, with AECs (active electrical cables) connecting them together.
Let’s get creative
AWS has a rich history of designing its own data center hardware. It previously teamed up with Broadcom to build custom network switches. In addition, its recently unveiled EC2 offerings feature its in-house networking stack, the Elastic Fabric Adapter (EFAv4), powered by its custom Nitro controllers, which optimizes network packet processing and reduces latency for high-throughput applications.
Each of the hyperscaler’s facilities contains what it calls “bricks” – a rack of AWS network switches designed to connect several servers to form more powerful clusters, or even other AWS data centers together in a scale-across fashion.
According to SemiAnalysis, AWS’s decision to continue down the custom route was due to the hyperscaler’s belief that Nvidia’s ConnectX-8 RoCEv2 NICs were “subpar” to their own NICs.
The complexity in its attempts to deploy its own custom NICs with the GB300 lies in its support for Elastic Fabric Adapter (EFA) compared to RoCEv2 Ethernet.
SemiAnalysis said it was “not convinced” that EFA-based NICs were better than RoCEv2 Ethernet on performance or user experience.
AWS’s insistence to shift away from Nvidia’s networking stack comes as the chip giant is dominating the infrastructure landscape. The hyperscaler’s decision to focus on employing its own custom solutions would help it avoid being locked into Nvidia’s ecosystem.
In the SemiAnalysis point of view, AWS’s GB300 rework helps to remove a single point of failure in Nvidia’s reference design, where each GPU talks to only one ConnectX-8 NIC.
“In AWS GB300 NVL72 design, each GPU talks to two K2v6 NICs, allowing workloads to not crash if one NIC fails,” the analyst firm wrote. “AWS bigly believes that EFA will be the future. Earth will see if AWS's big bet goes well.”
Despite Amazon’s insistence on its own networking components, Nvidia’s Blackwell-based rack-scale came out on top of a recent SemiAnalysis benchmark examining AI inference performance.
GB200 server systems using the chip giant’s original reference design reported the strongest performance across metrics, including throughput-per-dollar and tokens-per-megawatt, outpacing rival systems like the AMD MI355X.
Comments