Nvidia threw its hat in the CPU market today, unveiling its first Arm-based data center processor called Grace during the GTC 2021 keynote.

“Grace is a breakthrough CPU purpose built for accelerated computing applications of giant scale for AI and HPC,” Paresh Kharya, senior director of accelerated computing at Nvidia, said in a press briefing prior to the event.

Named for Grace Hopper, the U.S. computer programming pioneer, Nvidia’s first CPU is based on Arm Holding’s recently announced Neoverse N2 core architecture and will utilize a new memory subsystem and unified cache architecture to dramatically accelerate high-performance compute (HPC) environments like artificial intelligence (AI) training. Nvidia is currently seeking regulatory approval to acquire Arm in a deal valued at $40 billion.

Despite referring to Grace as a server-class CPU, the chip won't actually compete against processors from Intel, AMD, or even other Arm-based chips like Amazon’s Graviton2 or Ampere’s Altra. Instead, Nvidia claims Grace was designed from the start with large-scale HPC and AI workloads in mind.

That’s not to say Grace will be a slouch. Nvidia claims that when the chip launches in 2023 it will be capable of delivering a score greater than 300 in SPECrate 2017 integer base, a popular compute benchmark commonly quoted by chipmakers. If true, this would put Grace in line with the Ice Lake CPUs that Intel announced last week.

Nvidia Aims to Terminate AI Bottlenecks

Grace’s true purpose is to eliminate memory bottlenecks throttling the performance of larger AI workloads, according to Nvidia.

At current rates of growth, Kharya said AI models will reach 100 trillion parameters within the next few years, which presents a problem for current CPU architectures.

“Today’s architectures present serious bottlenecks in scaling the computing for these giant models,” he said. “GPUs have incredibly high compute rates and very high memory bandwidth… however, these giant models are so large that they cannot fit into the total GPU memory. We can share them in the system memory, but accessing them from the system memory is slow.”

In a four GPU configuration the aggregate memory bandwidth of the GPUs is about 64 Tb/s but if the GPU needs to leverage system memory that gets bottlenecked to just 512 Gb/s by the CPU, he explained.

Nvidia’s Grace CPU is designed to address these bottlenecks through three key features, Kharya said.

“First Grace uses next-generation Nvidia NVLink, which provides the fastest interconnect in the processor world. Each CPU is connected to a GPU with an incredible 900 gigabytes per second of bi-directional bandwidth,” he said. “Second, Grace has a new memory subsystem, leveraging [low-power DDR5] memory technology… It is twice the bandwidth of today’s DDR4, and is 10-times more energy efficient.”

The final feature is the use of Arm’s next-generation data center architecture: Neoverse N2. These three features together allowed Nvidia to eliminate memory bottlenecks hurting performance in high-density HPC workloads.

“The GPU can now access the CPU memory as fast as the CPU itself,” Kharya boasted. “With Grace, the AI community will have an optimal architecture to achieve peak performance for trillion parameter-plus models. These giant trillion parameter models, that would otherwise take months to train depending on the size of the cluster, can now be trained in just days.”

And for AI inference, Nvidia claims that Grace-equipped supercomputers will be able to process workloads in real time.

First Deployments of Nvidia's Grace CPU

Grace will see its debut in 2023 alongside a yet-to-be-announced Nvidia GPU in two new supercomputing clusters.

The first is the Swiss National Computing Centre’s Alps supercomputer which will be built in collaboration with Hewlett Packard Enterprise (HPE) and will replace the existing Piz Daint supercomputer. Once fully operational, the Cray EX-based system will serve as a general-purpose compute cluster available to a broad community of researchers.

“Alps will use the HPE Cray EX supercomputing infrastructure based on a cloud-native software architecture to implement a software-defined research infrastructure, as well as Nvidia’s novel Grace CPU to converge AI technologies and classic supercomputing in one single, powerful data center infrastructure,” Thomas Schulthess, a computational physicist at ETH Zurich and director at Swiss National Computing Centre, said in a statement.

Workloads planned for the system include climate and weather, materials sciences, astrophysics, computational fluid dynamics, life sciences, molecular dynamics, quantum chemistry, particle physics, and domains like economics and social sciences.

Nvidia claims Alps will be approximately 700% faster than its Selene supercomputing cluster, which went online last year.

Alps will come online around the same time as Los Alamos National Laboratory’s new Nvidia-based supercomputing cluster, which will also be built in collaboration with HPE.

“With an innovative balance of memory bandwidth and capacity, this next-generation system will shape our institution’s computing strategy,” Thom Mason, a director at Los Alamos National Laboratory, said in a statement. “Thanks to Nvidia’s new Grace CPU, we’ll be able to deliver advanced scientific research using high-fidelity 3D simulations and analytics with data sets that are larger than previously possible.”