AWS network engineer working on Project Rainier
– AWS

Amazon Web Services (AWS) recently switched the lights on for Project Rainier, a mammoth compute cluster powered by its custom Trainium2 chips.

The supercluster, to borrow hyperscaler parlance, contains nearly 500,000 of the home-developed hardware. It’s set to be used solely for Anthropic, the OpenAI rival taking on ChatGPT with its Claude foundation bot. Amazon has invested some $8 billion into Anthropic, becoming a minority stakeholder as well as the AI company’s chief cloud provider.

In terms of size, Rainier is on the Colossus end of the computing cluster scale, with a company blog post touting that it will house "tens-of-thousands" of the hyperscaler’s UltraServers spread across multiple mammoth data centers.

SDxCentral sat down with Ron Diamant, an AWS distinguished engineer and head architect of Trainium, to uncover the networking components behind one of the world's largest AI compute clusters.

NeuronLink & the scale-up domain

Rainier, first unveiled at AWS's Re:Invent event in Las Vegas in late 2024, is set to provide Anthropic with five-times more computing power compared to its prior largest training cluster. Each of its UltraServers features 64 Trainium2 chips, offering up to 83.2 FP8 petaflops of compute power.

AWS is connecting the hardware using NeuronLinks to support scale-up (server-to-server) while its Elastic Fabric Adapter (EFA) networking technology connects UltraServers across Rainier’s disparate data centers (scale-out).

On the scale-up side, each UltraServer consists of two racks positioned next to each other, with two servers per rack, all connected via the NeuronLink chip-to-chip interconnect – creating a scale-up domain of 64 total devices.

Trainium2 instance
Trainium2: AWS's custom-designed AI chip – AWS

The proprietary protocol runs on copper wires, offering what Diamant described as cost-efficient state-of-the-art bandwidth capabilities, two terabytes per second at approximately one microsecond per hop latency.

Diamant told SDxCentral that Rainier represented the first time the hyperscaler employed its NeuronLink-v2 technology beyond a single server. The ever-increasing intensity of AI workloads was cited as the reason behind the implementation.

“We saw that as AI models are becoming larger, some models actually require more than a single server to do inference,” Diamant explained. “That immediately led us to think through scale-up architectures. So that's the reason that we went and built the UltraServers.

“Each rack has two servers, and then we take two racks, put them next to each other, and they're all connected with NeuronLink," Diamant added. "With that technology, I'd say that the vast majority of models we're serving fit within a single UltraServer, and that means that we can make use of this dedicated scale-up technology in order to do fast inference communication, so things like tensor parallelism and so on.”

EFA: Multipathing at massive scale

But Rainier isn’t just a bunch of UltraServers connected in one single white space; it’s a series of massive interconnected facilities – aptly named, considering the project was christened after the 14,410-foot stratovolcano that can be seen from Amazon’s home of Seattle.

On the scale-out side, another bit of custom AWS networking technology does the heavy lifting: Elastic Fabric Adapter or EFA.

EFA is built on an in-house developed multipathing fabric protocol. Its key differentiator, according to Diamant, is its real-time path optimization capabilities. The system continuously tracks latency between any source and destination pairs, and if latency spikes even slightly, it immediately reroutes traffic along a different path.

“At a very large scale, every once in a while, a switch would crash or become congested, or anything like that," Diamant said. "Our ability to react to that at the millisecond scale and just vary the communication path is really critical.”

Project Rainier exterior
EFA manages the scale-out connections across the mammoth Rainier facility – AWS

Powering Rainier’s scale-out architecture is the third generation of EFA.

The technology helps to provide interconnections for what the hyperscaler externally refers to as UltraClusters. Internally, however, AWS engineers use a more technical designation: 10P10U, a name that reveals the network's specifications.

“It's tens of petabytes per second of bisection bandwidth across the UltraCluster, plus, on top of that, between any source and destination pair, you have 10 microseconds of latency,” Diamant said.

While those networking speeds sound impressive, Diamant confirmed the hyperscaler is already working on the next generation of EFA in order to “drive even lower latencies, and even more ability to react to hotspots, and very quickly bypass any congestion in the fabric.”

The multipathing component of the current Gen 3 iteration fundamentally changes how AWS can structure communications across Rainier.

While traditional fabrics work well with structured communication patterns like rings, where data moves sequentially from one node to the next, in Diamant’s view, they don’t scale efficiently.

“You send data from one node to the next, and as long as the underlying specs are solid, you'll get relatively good performance,” Diamant explained. “As you scale your training run or your communication world size to a bigger and bigger ring, the number of steps that you need to take just grows linearly, and performance starts to degrade.

“What we can do with EFA, due to this multipathing approach, is that instead of doing rings, we can send data from anyone to any source, to any destination, without creating hotspots on the fabric,” Diamant said. “So now we can scale with much better topologies, which allows us to scale to very large communication world size or giant clusters efficiently.”

EFA's multipathing also addresses challenges in modern inference workloads.

In disaggregated inference architectures, where prompt processing and output generation run on separate machines, traditional fabrics can suffer from problems that cause packet drops and latency spikes.

“What EFA gives us is the ability to identify these hotspots and find a different path to send the data to a different decode machine,” Diamant said, effectively enabling engineers to dynamically redirect Rainier’s traffic before congestion impacts performance.

Beyond its performance advantages, EFA includes built-in resilience mechanisms to handle the inevitable failures, an added bonus given the sheer size of Rainier, exponentially increasing the likelihood of something eventually failing somewhere throughout the stack.

“We definitely can route around failed switches and failed links,” Diamant confirmed. He also outlined to SDxCentral that the system can work around failing accelerators, with software optimizations able to quickly identify failed machines to reconfigure communication topologies.

What's next: Trainium3 & EFA Gen 4

While the lights went on at Project Rainier this past week, the story of Amazon building these mammoth AI clusters isn’t quite over.

In addition to Diamant’s comments on a fourth generation of EFA, AWS engineers are making strides on the hyperscaler’s next iteration of its custom hardware: Trainium3. Announced alongside Rainier at Re: Invent 2024, details on the next generation of AWS’s own accelerators are sparse at the time of writing, though prior reports suggest the chip will likely consume around 1-kilowatt of power.

With Rainier now up and running, AWS described the site as “a template” for deploying the computational power needed to power intense AI workloads addressing breakthroughs across everything from medicine to climate science.

Any future Rainier successor will likely feature even more powerful hardware with interconnect capabilities that’ll push latency boundaries even further.

But for now, Diamant and the team at AWS helped drive a mountain-sized project over the line by bringing together minds from across the business.

“On the one hand, we brought in Annapurna Labs with all the core technology, chip development, EFA, and so on,” Diamant said. “On the other, AWS has decades of expertise in building data centers and controlling the entire manufacturing line, so there are no third parties, and there's no long hops between one company to the other.

“We take the chips from TSMC (Taiwan Semiconductor Manufacturing Co.) and control the entire manufacturing line, land the chips in the data center, and qualify them at light speed," Diamant added. "That, to me, is the biggest thing that we did here. We launched a gigantic data center, maybe the biggest one in the world, in record time.”