Spectrum-x SuperNICs and Switch
– Nvidia

A little over a year ago, Nvidia CEO Jensen Huang took the stage at Computex 2024, the biggest semiconductor show in the world. Eyes from around the world were, as always, fixed on the man in the leather jacket and on what he might unveil to push the industry beyond Blackwell.

The headlines that day went to Rubin, the successor to Nvidia’s flagship Blackwell GPU, along with Huang’s announcement that the company would release a new high-end chip every year until 2027.

But while GPUs took much of the limelight that day, the most consequential announcement for AI’s future wasn’t a GPU at all.

Buried in the Nvidia CEO’s stacked keynote was Spectrum-X, an Ethernet-based networking fabric purpose-built for AI.

A year later, Spectrum-X has quietly become the plumbing that allows generative AI workloads to scale, re-engineering Ethernet to eliminate bottlenecks and overcome what Huang described last year as a history of being “designed for high average throughput."

SDxCentral sat down with the architects behind Spectrum-X to explore how Nvidia’s accelerated networking platform has evolved in its first year, and what it means for the next wave of AI infrastructure.

But first, what is Spectrum-X?

Nvidia Spectrum-X
– Nvidia

Spectrum-X is an end-to-end accelerated networking platform designed to improve the performance and efficiency of Ethernet-based AI clusters.

The offering combines Nvidia’s networking hardware and software to power a networking platform specifically optimized for AI workloads. It brings together the Spectrum SN5600 switches, part of Nvidia’s Spectrum-4 switch architecture, with Nvidia’s BlueField 3 data processing units (DPUs) to boost generative AI network performance by 1.6-times over traditional Ethernet fabrics.

Spectrum-X also integrates network-level remote direct memory access (RDMA), which helps provide that performance boost, along with adaptive routing to significantly reduce packet congestion.

Foundation behind the fabric

The roots of what would eventually become Spectrum-X can be traced back to 2019, when Nvidia beat the likes of Intel, Microsoft, and Xilinx to acquire Mellanox in a deal worth $6.9 billion.

The Israeli-American firm was a leading developer of interconnect networking technologies for data centers and high-performance computing (HPC). Having built a host of Ethernet-based switches prior to the purchase, many of the minds behind Mellanox’s offerings would later use their expertise to help craft Spectrum-X

One of those minds was Kevin Deierling, Nvidia’s SVP of networking, who said that Huang was the only industry leader who understood that the data center was the new unit of computing and not just individual servers with GPUs.

“I was looking at a picture of [Huang] when he first came to Israel back in 2019 after Nvidia had won the bidding war, and remembering his comments that people thought of an accelerated computer as a box that has an x86 in it, with sheet metal wrapped around it for memory," Deierling said. “At that time, maybe 5% of the servers might have a GPU in them. But Huang said that's actually not the computer; the data center is the new unit of computing. And if you take that as your starting point, that’s the vision.”

That data center-centric vision required rethinking networking from the ground up, setting the stage for what would become Spectrum-X.

Deierling noted that key technologies powering Spectrum-X, like RDMA over Converged Ethernet (RoCE), stemmed from Mellanox's “historical legacy” in what would later be reworked by Nvidia to optimize AI workloads.

Breaking bottlenecks: Why traditional Ethernet fails for AI

While Spectrum-X uses Ethernet, the technology hasn’t always been the core for AI networking. Traditional Ethernet was built for throughput and multinode communication across a data center or even the entire internet. But for AI, Ethernet would suffer from jitter problems, making GPUs running intense workloads finish at different times and subsequently having a negative impact on outputs.

"If you have 100,000 GPUs and 99,999 finish at the same time but just one is late, that one GPU determines the performance of the entire data center because everyone is waiting for data from that GPU," Gilad Shainer, a networking SVP at Nvidia, explained.

The root cause of Ethernet’s issue lies within what's known as Amdahl's Law. While operators can parallelize most AI computation across thousands of GPUs, there will always be a serialized portion where all results must be gathered and processed collectively.

Deierling likens the issue to pizza, in that if a store needs to make more pizzas, adding more ovens might cook more pies, but it doesn't help if each pizza still takes 30 minutes to cook and subsequently deliver.

Domino's Pizza meme
An accurate depiction of Deierling's pizza analogy (gone wrong) – Domino's Pizza

When congestion occurs – an inevitability in distributed computing – packet drops must be retransmitted, creating the exact synchronization issues that cripple AI performance. This creates what Shainer calls "noise between workloads," where one job running across multiple servers impacts another job sharing the same network infrastructure.

Nvidia’s solution was to copy what works for InfiniBand and bring it to Ethernet.

The chip giant’s engineers took three core innovations from its traditional InfiniBand and applied them to Ethernet: implementing lossless networking to eliminate retransmission delays; creating adaptive routing that spreads traffic based on both local and global network conditions; and developing telemetry-based congestion control that prevents jobs from interfering with each other.

“We copied as much as we could from InfiniBand, and Spectrum-X is the closest you can get to InfiniBand,” Shainer explained. “We didn't build a switch. We built an infrastructure. Part of the functions are in the switch. Part of the functions are in the superNIC, and they work together to spread the traffic not just according to its local situation.”

To bring the pizza analogy back, Deierling likened it to putting a pizza oven in the delivery truck: "We're processing the data as it's moving through the network.”

Nvidia’s team gave Ethernet the tools to challenge InfiniBand's dominance in AI networking, while leveraging the existing software ecosystems that enterprises had built around Ethernet infrastructure.

That innovation positioned Ethernet to challenge InfiniBand's crown as the AI back-end leader. As recently as 2023, InfiniBand held an 80% share of the AI back-end networks market. But recent Dell’Oro Group research suggests that increased interest in Ethernet-based offerings could drive nearly $80 billion in data center switch sales over the next five years.

That report suggested networking engineers were far more familiar with Ethernet compared to InfiniBand, and that the hefty costs required to interconnect hundreds of thousands of GPUs meant Ethernet would eventually win out over its rival.

So does Spectrum-X mean Nvidia has abandoned its long-standing love for InfiniBand? No. Far from it, in fact.

Alongside Spectrum-X, an InfiniBand-based alternative was also unveiled at Computex ‘24, dubbed "Quantum-X," enabling the chip giant to offer end-to-end networking platforms tailored to customer needs.

“Not every data center can handle InfiniBand, because they've already invested their ecosystem in Ethernet for too long, so what we've done is we've brought the capabilities of InfiniBand to the Ethernet architecture, which is incredibly hard," Huang explained during that Computex keynote.

Co-packaged optics: The ‘next level of infrastructure’

Nvidia Spectrum-X photonics
– Nvidia

In the year following Spectrum-X’s unveiling, the Ethernet-based networking platform has already found itself powering some of the largest and most intense AI workloads in the world.

The platform has been in use in Colossus, the supercomputing cluster from Elon Musk’s xAI. It’s also set to feature in the first few Stargate data centers, the $500 billion digital infrastructure project from OpenAI, Oracle, SoftBank, and Abu Dhabi's MGX.

But the biggest update to Spectrum-X in its first year came from yet another industry event: Nvidia’s own GTC back in March. There, the chip giant revealed that both Spectrum-X and Quantum-X would be gaining support for silicon photonics switches.

Compared to what the company described as “traditional methods,” the newly added co-packaged optics (CPO) would offer 3.5-times more power efficiency, 63-times greater signal integrity, and 10-times better network resiliency at scale. At the time, though, Nvidia failed to disclose exactly what it was comparing its CPO-enhanced Spectrum-X to.

Deierling told SDxCentral that back in 2021, he took a “clobbering” from Nvidia’s optics team for saying CPO was a great technology that was five years away.

But four years later, the addition of CPO support adds to Nvidia's broader vision of optimizing across the entire data center stack, with the technology adding to improvements alongside congestion control, automatic path migration, and packet scattering

Shainer likened CPOs to “the next level of infrastructure.”

An inside look at the CPOs in Nvidia's Spectrum-X
An inside look components in Nvidia's networking platform, including new co-packaged optics – Nvidia

Unlike single-server workloads that primarily use copper cabling, AI supercomputers utilize a multirail topology with a huge amount of optical fiber. Shainer explained that optics alone can account for up to 10% of total compute power, adding substantial energy demands to large-scale AI data centers. The transceivers, meanwhile, require frequent replacement, with human intervention during maintenance a leading cause of downtime.

By switching to CPOs and integrating the optical engine directly with the switch package, the Shainer outlined that it eliminates the distance between the transceiver and switch, leading to reduced power consumption and maintenance, as well as reduced overall costs.

While CPO isn't a new concept, Shainer suggested that previous attempts at production failed due to oversized optical engines and difficulties in high-yield packaging.

“One [issue] is that the optical engines that were created were too big. So far, it was based on Mach-Zehnder architecture," Shainer said. "The second thing is that people could not find a way to package CPO in a way that you can validate that in a high yield, which means that it's too expensive to actually build."

Nvidia addressed these issues by developing a new optical engine architecture using micro-ring modulators.

“These are very small, but the challenge is that they can create heat and you need to control it," Shainer explained. "This is where two attempts to do microwave modulators failed before Nvidia started to do the work ... but we managed to build a micro modulator that supports large switches.”

The designer then worked with partners like TSMC to create a new packaging architecture called Compact Universal Photonic Engine (COUPE), which brings together a 65-nanometer electronic integrated circuit (EIC) and a photonic integrated circuit (PIC), adding transceivers capable of supporting a multitude of distances.

“When you go to CPO co-packaged optics, you’re going to pick one [transceiver] … what we have done in our CPO is cover short distances, and long distances in and between data centers," Shainer said. "That’s where the magic is.”

Spectrum-X: One year on

A little over a year since its unveiling, Spectrum-X has quietly become the foundational plumbing for some of the world's most demanding AI workloads. But more telling than high-profile deployments like in Colossus is what Nvidia has learned about AI networking in Spectrum-X's first year of operation. As Deierling puts it, the company discovered that “raw speeds and feeds aren't enough.”

The lessons have been hard-earned. Nvidia builds its massive supercomputers internally to develop its own AI models and test open-source ones, which gives the chip giant what Deierling describes as "fine-grained visibility" into where bottlenecks occur.

"If you look at something like a training job that runs for three months, my goodness, how do you speed that up? You really need an in-depth understanding of AI models and training,” Deierling said. “Even a year and a half ago, internally, people were saying, 'oh, we're just going to run that on a single node. We don't really care about the network.' It turns out that none of that is true, with inference becoming more complicated with different models and questions in an agentic flow, mixture-of-experts, and different data that needs to be protected between users.”

Shainer outlined that the AI-focused networking platform was built for what he described as a "zero-revenue market" at the time, but one that’s only going to explode in the coming years the technology continues to proliferate.

The question isn't whether Spectrum-X itself will continue to evolve – with co-packaged optics integrations set to arrive in 2026 – that's a given. The real question is how the broader industry will adapt to the vision Nvidia’s founder outlined back in 2019 to make the data center a fundamental unit of AI computing, with networking as its critical backplane.