Generic hardware/semiconductor image
– Getty Images

One can argue the excitement has been taken out of the hardware space, what with the big boys insisting on producing annual roadmaps akin to how Marvel maps out their upcoming movies. But there’s more to high-performance hardware than just GPUs, and in that vein, 2025 did not disappoint.

From some truly staggering custom hyperscaler hardware to mega-deals between some of the biggest names in this space, innovation and scale show no signs of slowing. And yet, just on the horizon, a potential bump in the road looks to be getting larger in the worrying memory malaise, seeing firms scramble to snap up modules before prices soar (even more than they already have done).

Here’s a snapshot of all the major hardware-related updates from the past 12 months.

OpenAI shacks up with Broadcom

Broadcom sign
– Getty Images

Arguably one of the biggest deals of the year, OpenAI’s tie-up with Broadcom marked the culmination of years of wanting for the ChatGPT maker to have its own chips. Now, with Broadcom, it’s going to get that, with Broadcom aiding in the joint development of its own GPUs.

OpenAI looks to have made a sensible decision given Broadcom’s history of working with big-name firms developing custom hardware. It previously supported the creation of Google’s Tensor Processing Units (TPUs) – another tick, given OpenAI’s hardware team is made up of engineers who worked on Google’s TPU projects.

OpenAI’s Broadcom deal summed a year in which the former moved away from its historic ties to Nvidia. While the pair signed an agreement earlier this year, it’s since moved to diversify its chip supply, culminating in an agreement with bitter rivals AMD.

OpenAI's multipronged chip strategy, now supported by its own hardware built by Broadcom, essentially echoes what its hyperscale colleagues did some years ago: enjoy long portions of a particular supplier, then dip its toes in to rivals, before seeking out its own offerings, ala Google with its TPUs and Amazon Web Services' (AWS) Trainium.

Read the full story

OpenAI enlists Broadcom to co-develop custom AI hardware

OpenAI's Broadcom deal signals the end of Nvidia's AI chip monopoly


AWS doubles down on custom silicon

Trainium3 ultraservers rack
– AWS

Sticking with custom hardware, AWS went one step further, unveiling a mammoth compute cluster powered by its own hardware, while also further building out its custom stack.

In November, it revealed Rainer to the world, in what is a supercluster (to borrow Amazon parlance) containing some 500,000 of its home-developed Trainium2 chips.

Claude developer Anthropic is the lucky recipient of Rainer, using it to train its next-generation models and power inference workloads. Amazon does, of course, own a minority stake in Anthropic.

Each of the UltraServers Anthropic will have access to inside of Rainer contains 64 Trainium2 chips, offering up to 83.2 FP8 petaflops of compute power. SDxCentral got the inside scoop on the hardware AWS is using to connect all of those together: its own NeuronLink-v2 tech to support scale-up (server-to-server).

AWS' custom love-in didn’t stop inside the white space, either, with the hyperscaler’s Elastic Fabric Adapter (EFA) networking technology connecting UltraServers across Rainier’s disparate data centers (scale-out).

EFA is built on an in-house developed multipathing fabric protocol, and is capable of continuously tracking latency between any source and destination pairs. If latency spikes even slightly, it immediately reroutes traffic along a different path.

To further add to its already strong custom stack, AWS showcased its latest generation of Trainium – Trainium3 – in December. Marking the company’s first foray into 3-nanometer, AWS claims UltraServers featuring Trainium3s offer up to 4.4-times more compute performance, four-times greater energy efficiency, and nearly four-times more memory bandwidth than the previous generation.

And to top off its internal tech efforts, the hyperscaler introduced custom dense wavelength-division multiplexing (DWDM) transponders for both metro and long-haul connections, AWS claimed offer 73% more bandwidth than compared to its initial attempt at a transponder.

Read more

The network engine behind AWS’ massive Rainier supercluster

AWS makes Trainium3 UltraServers generally available

AWS expands custom silicon push with homegrown fiber optic transponders


The triumphant return of Nvidia’s BlueField line

Nvidia Bluefield-4
– Nvidia

In 2021, Nvidia’s BlueField-4 data processing unit (DPU) was teased for the first time. But after that … nothing. Nadda. Zilch.

But the October 2025 edition of its GTC roadshow saw the processor in all its glory. And, in line with Nvidia’s wider stack, it received an AI-focused overhaul – with Jensen Huang touting it as “purpose-built for the gigascale AI era.”

In what was the first industry deep-dive into BlueField-4, SDxCentral secured a conversational coup with Itay Ozery, Nvidia’s director of product marketing for networking, who likened BlueField-4 to the “processor for running the operation system of the AI factory.”

“What we've done since 2021 is we realized that, hey, we need to optimize the AI compute fabric so it can support the scaling laws of AI with the Hopper and Blackwell generation needing to scale to tens-of-thousands of GPUs and even hundreds-of-thousands of GPUs,” Ozery said following its launch.

And with other changes to it, including ensuring alignment with Nvidia’s Ethernet-based Spectrum-X networking fabric, the DPU might not be the flashiest name in the company’s hardware roster, but BlueField-4 looks to be providing sufficient headroom for GPU-to-GPU communications ahead of the shift to the next generation Vera Rubin platform.

For more on Nvidia and BlueField-4

Nvidia's BlueField-4: A first look at the DPU built to run AI factories

Nvidia reveals next-gen DPU to help offload gigascale AI infrastructure

Inside Spectrum-X: Nvidia’s Ethernet networking platform


Memory malaise begins to set in

Server memory
– Thinkstock / NorGal

A potentially worrying trend appears to be gaining traction as the year draws to a close: the demand from AI developers for memory hardware could lead to a component shortage.

Hewlett Packard (HP) warned that supply chain issues are a genuine roadblock set to impact its profitability in 2025. Dell and Lenovo have warned of cost increases for memory and storage components and are mulling price increases. And Micron even closed up shop entirely on its consumer division in a bid to focus solely on enterprise AI demand.

SK Hynix has drawn up plans for a low-power wafer factory to meet spiraling demand, while Samsung is reportedly halting production of its SATA SSD line, which, like all solid-state drives, relies on that other hot commodity in the AI race, NAND flash memory.

Already at a critical juncture, there are suggestions – not concrete ones, but reports – that Nvidia’s Rubin line of GPUs may also be impacted as SK Hynix has opted to shift the ramp-up of its HBM4 chip offering in the third quarter of next year, instead of the second quarter as previously planned.

With Nvidia now teaming up with SK Hynix to create SSDs specifically designed for AI, the big boys appear to be shoring up their supply chain with something in the works. Further, suggestions that Nvidia has moved to snap up significant capacity of electro-absorption modulated lasers (EML), we could see a potential component shortage sometime in 2026 or 2027.

Read more

Nvidia and SK Hynix partner on AI inference chips

Nvidia's aggressive laser procurement spurs supply chain fears

SK Hynix predicts AI memory gloom until 2028

Samsung riding AI memory wave to top spot as server prices jump


Google’s staggering next-gen TPU

Google Ironwood TPU
– Google

Like AWS, Google’s custom silicon efforts continued to soar in 2025 with the announcement of its seventh-generation Tensor Processing Unit (TPU), codenamed Ironwood. Unveiled in April and teased in November, it’s specifically engineered for inference, with each individual chip delivering a staggering peak compute of around 4,600 teraflops.

The true scale of Ironwood lies in its ability to interconnect up to 9,216 chips into a single pod, creating a system capable of a staggering 42.5 exaflops. This is achieved using Google’s proprietary networking tech, Inter-Chip Interconnect (ICI), which provides 400 Gb/s (400G) bandwidth in each direction, effectively creating one massive "brain" to support fast-moving workloads.

A key context point for the compute figures, however, is the precision. While the 42.5 exaflops measurement is calculated using FP8 (8-bit) precision, traditional supercomputing benchmarks like those for El Capitan are in FP64 (double) precision.

For more on Ironwood

Google claims its next-gen TPUs offer ‘more compute’ than the world’s fastest supercomputer