Close up of Arm's new AGI CPU
– Arm Holdings

The enterprise CPU market is thriving. Just a few years ago, all eyes were on the shiny, high-power GPUs. But the surge in demand for agentic AI, with its need for complex logic, combined with the rising tide of the memory shortage resulting in more inactive data being offloaded to the processor, means we’re truly witnessing a renaissance.

While Nvidia has sought to make strides in a market dominated by Intel – despite its financial woes – and AMD, a new challenger approaches – one very familiar to this space.

At its annual Everywhere event in San Francisco, Arm shocked the world with news that it was extending its design expertise into production silicon products for the first time in the company’s history.

Having licensed its architecture for the longest time, the company unveiled its own processor for AI data centers, the Arm AGI CPU, which it claims offers more than 2-times the performance per rack compared with x86 platforms.

“AI has fundamentally redefined how computing is built and deployed. Agentic computing is accelerating that change,” Arm CEO Rene Haas said. “Today marks the next phase of the Arm compute platform and a defining moment for our company. With the expansion into delivering production silicon with our Arm AGI CPU, we are giving partners more choices all built on Arm’s foundation of high-performance, power-efficient computing, to support agentic AI infrastructure at global scale.”

Rene Haas, CEO, Arm holding aloft the new Arm AGI CPU
Rene Haas, CEO, Arm holding aloft the new Arm AGI CPU – Arm Holdings

In spec terms, the AGI CPU is built on Taiwan Semiconductor Manufacturing Company's (TSMC’s) three-nanometer (3nm) process and features more than 130 Neoverse V3 cores. Some 12 channels of double data-rate five (DDR5) memory provide more than 800 Gb/s of aggregate memory bandwidth.

The chip is front and center of Arm’s 1OU Dual Node modular rack server reference architecture, encompassing deployments of up to 8,160 cores per standard 36-kilowatt (kW) air-cooled rack.

Upon unveiling its hardware highlight, Arm said more than 50 firms were helping bring its compute platform into silicon, with Meta set to be the launch partner. The pair signed a multi-year partnership last October, with Arm’s Neoverse designs already leveraged in Meta’s data center platforms to power its AI ranking and recommendation systems across the social media firm’s family of apps, including Facebook and Instagram.

But this latest move goes one step further. Meta is the codeveloper of the AGI CPU, with it specifically optimized to work alongside its aptly named Meta Training and Inference Accelerator (MTIA) custom silicon.

“We worked alongside Arm to develop the Arm AGI CPU to deploy an efficient compute platform that significantly improves our data center performance density and supports a multigeneration roadmap for our evolving AI systems,” Santosh Janardhan, head of infrastructure at Meta, explained.

Agentic AI and memory shortages drive the CPU renaissance

Full production availability is anticipated in the second half of this year, with the likes of Lenovo, Supermicro, and ASRock Rack enlisted to bring early systems to market.

But as Arm gears up to step into a brave new world after 35 years of licensing its IP, it begs the question: why now?

The answer lies in the fact that CPUs do what GPUs simply aren’t designed for. All the hype around agentic AI and its ability to perform complex tasks means related workloads require a heck of a lot more orchestration than simple chatbot systems. Think of the GPU as a brute force machine, powering workloads at high speed, but with little to no "brains," for lack of a better term.

It’s why Nvidia’s Groq line or Cerebra’s wafer-scale systems are gaining traction of late; agentic AI requires more than what the GPU can offer.

The CPU renaissance looks to be part of that play, accounting for more latency-intensive needs while leaving the GPU to handle the model inference. And thanks to sizable memory bandwidth and input-output (I/O) in modern CPU architectures – including Arm’s new processors – CPUs can move related agentic data with limited lag.

And with the added headache of the memory shortage, CPUs are being paired more closely with GPUs to create more unified memory pools through systems like the Compute Express Link (CXL) acting almost like an extension of local RAM to offload "lesser" or inactive data when running longer models.

All of these considerations have led to the resurgence of interest in the CPU – along with the wider stack beyond just the GPU – as operators look to push the demand of their entire offering to the limit.

Nvidia's Vera CPU
Nvidia's Vera CPU – Nvidia

Nvidia's CPU train has left the station as Arm jumps on board

Arm isn’t alone in vying for a piece of the increasingly more lucrative CPU pie. Intel and AMD’s x86 architecture dominate the server space, with the pair having previously formed an ecosystem advisory group to fend off rising competition from Arm.

But a newer player is looking to turn the market on its head: Nvidia.

The long-time GPU player stepped into the CPU space in early 2021 with its Grace line and is currently in full production of its Vera data center CPU, which is expected to be available in the second half of this year.

The likes of Nebius, Nscale, Oracle, and Alibaba have all lined up to offer Vera to cloud customers, while Cisco, Dell Technologies, and Hewlett Packard Enterprise (HPE) are set to support the CPU in server lines.

Vera encompasses 88 custom-designed "Olympus" cores, each of which can run two tasks. Meanwhile, Nvidia has opted for low-power double data-rate 5X (LPDDR5X) memory, which runs at lower power than DDR5 while offering up to 1.2 Tb/s of bandwidth.

Speaking to SDxCentral at the recent GTC event, Shar Narasimhan, Nvidia’s director of product marketing for AI and data center GPUs, said its push into this space was due to CPUs becoming as important in accelerating computing as the GPU.

“We've spent a lot of time optimizing the GPU," Narasimhan said. "We've seen the bottlenecks that are present when it's waiting for instructions and so we've realized not only do we need to get the data faster from the CPU, because it is the lead orchestrator in the data center, but we also need to make sure that when the CPU has to do its verification role it is now a sort of harmonious loop between the two.”

Narasimhan explained the potential boost with an example from the world of reinforcement learning:

"The GPU does an iteration, and then it has to go complete testing. You have validation scripts that are run. I make a certain hypothesis or an inference, I now need to go test that to make sure I made the correct prediction. And that response has to come back to the GPU, and that gives the GPU accurate information to do its next inference. Those inferences are like a train leaving the station. They just keep going. You don't want to have the GPU wait for that response to come back because now you're not utilizing a very precious asset.

“The other thing you also don't want to do is have the GPU train on inaccurate information, because the train left the station before that response came back. So by making the CPU just as fast – so that that loop can produce proceed continuously – that's how we continue to accelerate data centers.”