Intel's next-generation Xeon Scalable processor family, which it’s calling Sapphire Rapids, is shaping up to be a turning point for the embattled chipmaker despite delays pushing its release back to early next year. Late last month, the company shed some light on the upcoming data center chip during the annual Hot Chips 2021 conference.

Sapphire Rapids is a mile marker of sorts for Intel. It will be Intel’s first to fully embrace a chiplet architecture — Intel calls these tiles — and it will be the first mainstream data center processor that supports DDR5, high-bandwidth memory, PCIe Gen. 5.0, and compute express link (CXL).

“Sapphire Rapids delivers a step function in performance across a broad set of scalar and power workflows,” Arijit Biswas, principal engineer at Intel, said during a presentation at Hot Chips 2021.

Intel Embraces Chiplets

Sapphire Rapids will see Intel abandon monolithic dies, like those used in this year's Ice Lake Xeon Scalable, in favor of multiple compute tiles that are packaged together under an integrated heat spreader.

“At the heart of Sapphire Rapids is a modular, tiled architecture that allows us to scale the Xeon architecture beyond physical reticule limitations,” Biswas said.

These tiles are interconnected using Intel’s embedded multi-die interconnect bridge (EMIB) technology, which allows them to communicate with each other and share resources. Using the technology, “we are now able to increase core counts, caches, memory, and I/O,” Biswas said.

If this sounds familiar, that’s because AMD did the same thing four years ago with its EPYC, ThreadRipper, and later Ryzen processor families. AMD’s latest EPYC processors feature up to eight chiplets, each with up to eight cores, and 32 megabytes of level 3 (L3) cache for a total of 64 cores, 128 threads, and 256 megabytes L3 cache.

Chiplet architectures have a number of advantages over monolithic designs, which will likely help Intel compete against AMD. The modularity also allows chipmakers to increase core counts dramatically since adding cores can be achieved by adding more tiles to the processor package.

And thanks to improving interconnect technologies like AMD’s Infinity Fabric and Intel’s EMIB, chipmakers have been able to minimize latency challenges associated with die-to-die communications.

According to Biswas, Sapphire Rapid’s compute tiles will have full access to all resources — including cache, memory, and input/output (I/O) functionality — on all tiles. This means any one core will have access to all of the resources on the chip and are not limited to what’s built into the tile.

So while Intel is taking a cue from AMD, it appears Intel's chips won't face the same kind of cache limitations as seen with EPYC, which despite having as much as 256 megabytes of L3 cache, only 32 megabytes are available to a core.

A Faster, Smarter Core

The compute tiles themselves will use Intel’s 10-nanometer SuperFin process node, which it recently rebranded as Intel 7, and will feature a revamped performance core — codenamed Golden Cove — alongside a host of hardware accelerators.

“The new performance core in Sapphire Rapids brings significant scalar performance improvements,” Biswas said. “Additionally, multiple integrated accelerator engines and increased core counts provide for a massive improvement in data parallel performance.”

Intel claims the core is optimized for containerized, cloud-native, and artificial intelligence workloads, and will achieve a 19% instructions-per-clock uplift over this year’s Ice Lake processors.

Beyond higher raw performance, the core also features three new hardware accelerators. These include:

  • Accelerator Interfacing Architecture (AIA)
  • Advanced Matrix Extensions (AMX)
  • Data Stream Accelerator (DSA)

Intel’s AIA is tasked with facilitating the dispatch of workloads to the CPU’s hardware accelerators. Similarly, The DSA is responsible for offloading common data movement tasks, which would otherwise bog down the CPU cores.

Finally, Intel’s AMX promises a substantial performance uplift for tensor workloads like deep learning artificial intelligence algorithms — up to 700% higher performance compared to advanced vector extension (AVX) 512.

I/O Up the Wazoo

If higher core counts and a nearly 20% IPC uplift weren’t enough, Intel’s Sapphire Rapids boasts support for a menagerie of next-generation I/O and memory tech including PCIe Gen. 5.0, CXL 1.1, DDR5, and on-die HBM.

The latter is particularly notable as HBM, while common in GPU architectures, is relatively new when it comes to CPUs. Intel claims the onboard memory dramatically improves performance for memory-sensitive applications.

These technologies open the door to servers with larger capacities, faster and more flexible peripherals, as well as larger, more complex workloads. To this end, Sapphire Rapids is already slated to power the Argonne National Laboratory’s Aurora Supercomputer.

“Integrating high-bandwidth memory into Intel Xeon Scalable processors will significantly boost Aurora’s memory bandwidth and enable us to leverage the power of artificial intelligence and data analytics to perform advanced simulations and 3D modeling,” Rick Stevens, associate laboratory director of Computing Argonne National Laboratory, said in a statement.

Perhaps more importantly for Intel, it’s bringing these technologies to market roughly six months ahead of rival AMD. AMD isn’t slated to launch its EPYC 4 Genoa processors and Zen 4 architecture, which adds support for many of these technologies alongside a process-node shrink to 5 nanometers, until late 2022.

Staying On Track

Of course, all of this is dependent on Intel staying on track. Sapphire Rapids was already delayed by about six months earlier this year to provide additional validation time.

In a status update in July, Lisa Spelman, VP and GM of Intel’s Xeon and Memory Group, wrote that “given the breadth of enhancements in Sapphire Rapids, we are incorporating additional validation time prior to the production release.”

Had it not been for the delay, Intel’s roadmap had Sapphire Rapids slated for release in the fourth quarter of 2021.

However, the delay came as no surprise to analysts who hadn’t expected Sapphire Rapids to reach the market until spring 2022. In an earlier interview, Baron Fung, research director at Dell’Oro Group, said he anticipated higher demand for Sapphire Rapids compared to previous refreshes.

“It appears that customers are more likely to skip IceLake and wait for Sapphire Rapids instead,” he said. “Sapphire Rapids has some nice improvements over Ice Lake, notably CXL, PCIe 5.0, and DDR5, which I think any customers would want to wait for.”