AI modeling is demanding.

Particularly, large language models (LLMs) can eat up massive network bandwidth, slowing traffic and creating bottlenecks — ultimately reducing time-to-market in an ever-competitive race to deploy advanced AI tools.

Semiconductor and infrastructure software provider Broadcom says it can help solve this problem: The company today rolled out Jericho3-AI, which is designed to provide ethernet networking for AI at scale. The high-performance fabric offers enhanced load balancing, congestion-free operation, high radix and zero-impact failover, according to Ram Velaga, SVP and GM of Broadcom’s core switching group.

“The benchmark for AI networking is reducing the time and effort it takes to complete the training and inference of large-scale AI models,” he said. “Jericho3-AI delivers significant reduction in job completion time.”

Different traffic requires different AI tools

The new tool comes as AI adoption and AI modeling  accelerates by the day: According to a new forecast from IDC, global spending on AI will reach $154 billion this year — representing an increase of 26.9% over 2022 — and will reach $300 billion (or more) by 2026.

“Companies that are slow to adopt AI will be left behind — large and small,” according to Mike Glennon, senior market research analyst with IDC's customer insights and analysis team. “AI is best used in these companies to augment human abilities, automate repetitive tasks, provide personalized recommendations and make data-driven decisions with speed and accuracy.”

The big issue is that large data center operators — at Microsoft, Meta, Google, for example — are building very large clusters to train LLMs, “which are basically super computers,” explained Bob Wheeler, principal analyst at Wheeler’s Network.

For instance, OpenAI-built ChatGPT using 10,000 GPUs, and enterprises are continually investing in infrastructure consisting of clusters with tens of thousands of GPUs.

But, Wheeler explained, the network traffic generated when training AI models is very different from traffic for cloud computing instances. It has different flow lengths and sizes and is going to different places. Also, in AI training, there comes a juncture where all GPUs have to synchronize.

“At that point, you’re essentially stalling the training waiting for the last node to complete,” said Wheeler.

The node that takes the longest to complete is stalling the entire cluster. If there's congestion, or a packet gets dropped and needs to be re-transmitted, the model takes longer to train. “It’s very important to keep operations as fast as possible,” he said.

Congestion-free by design

Velaga agreed that AI has “turbocharged” the amount of processing power and data moved between accelerators on GPUs. AI workloads have a low number of large, long-lived flows that all start concurrently upon completion of an AI computation cycle, he explained.

In AI, if two GPUs communicate, they talk to each other for a long time, he explained. But if one of those GPUs wants to talk to a third GPU, they can get backed up waiting for the lane to clear even if other lanes are available. Ultimately, AI models require all GPUs to work together.

“As simple as it might sound, it is a very hard problem to solve,” said Velaga.

The Jericho3-AI fabric provides load balancing that equally sprays traffic over all links to help maximize network utilization. This allows for “congestion-free” operation with traffic scheduling to avoid “flow collisions” and jitter.

Velaga compared it to a highway system with four lanes. If all cars enter one lane and the remaining three are available, the vehicles can’t reach the speed limit of 65 miles per hour. This is because just a small percentage of cars slow all the others down.

With Jericho3-AI, when traffic is entering the proverbial highway, the system is aware that other traffic is getting off at exit C, so there is capacity there.

The fact that the fabric’s architecture eliminates congestion by design is a unique feature, Wheeler said. It can also be built out to a very large scale in a “very consistent and manageable way.”

“Jericho3-AI offers a high-bandwidth, low-latency and low-power choice for networks connecting tens of thousands of GPUs, revolutionizing the economics of building and maintaining AI clusters for this exciting new era,” he said.

Ethernet reigns supreme for AI modeling

Additionally, Jericho3-AI’s “ultra-high” radix allows the fabric to scale connectivity to 32,000 GPUs — each with 800 billions of bits per second — in a single cluster. And a zero-impact failover functionality provides automatic path convergence that doesn’t impact job completion time, Velaga said.

The fabric offers 26 petabits per second of ethernet bandwidth and delivers 40% lower power per gigabit and also includes long-reach SerDes, distributed buffering and advanced telemetry that provide flexibility in architecture and deployment.

According to Velaga, Jericho3-AI provides at least 10% shorter job completion times versus alternative networking tools — notably InfiniBand networking products offered by Nvidia, Intel, Oracle, IBM and others.

“Eventually the only technology that wins and will continue to win and will survive is ethernet,” said Velaga. “Nothing has been as pervasive, as commoditized and generally available with a broad ecosystem that drives innovation like ethernet.”