AMD unveiled its third-generation EPYC data center chips today, which the company claims are more than twice as performant as Intel’s highest-end CPUs.

“Third-gen EPYC is the highest performance server processor,” boasted Ram Peddibhotla, EPYC product manager at AMD.

Codenamed Milan, AMD’s EPYC 3 is largely an architectural update to the company’s flagship CPUs. Like the previous generation, EPYC 3 tops out at 64 cores/128 threads, is based on the same Taiwan Semiconductor Manufacturing Co. (TSMC) 7-nanometer process node, and is even drop-in compatible with servers designed for EPYC 2 via a BIOS update.

Despite the similarities between the two generations, AMD managed to deliver a 19% increase in instruction per clock over EPYC 2.

And in a veiled jab at the company’s biggest competitor, Intel, Peddibhotla said AMD had managed to achieve these performance gains on schedule and without delay. Intel, by comparison, has struggled to bring its 10-nanometer Icelake data center chips to market, having only just recently entered volume production.

An EPYC New Zen 3 Core

According to Mike Clark, corporate fellow and silicon design engineer at AMD, the performance gains achieved by EPYC 3 are thanks in large part to a redesigned core architecture called Zen 3, which launched with the company’s 5,000-series client CPUs last year. Zen 3 introduced several notable enhancements including an updated cache architecture and faster branch prediction. Improvements to the latter allowed AMD to significantly bolster EPYC 3’s floating-point integer performance by increasing the speed at which the branch predictor could load instructions into the pipeline, Clark explained.

One of the biggest changes to Zen 3, and consequently EPYC 3, is a revised core cache complex, which provides each processor access to a larger pool of level three (L3) cache. Clark explained that while the total available L3 cache remained unchanged at 256 megabytes, the move to Zen 3 increased the amount of L3 cache available to any one core from 16 megabytes to 32 megabytes.

“We now have 32 megabytes across eight cores, so those eight cores can talk more effectively to each other. ... In lower thread count scenarios, we have more L3 cache per thread, per core, and therefore reduce the effective latency to memory,” he said.

The change offers significant performance advantages for certain workloads like database applications with a large instruction footprint per thread, added Noah Beck, server system on chip architect at AMD.

In this scenario, the large shared cache allows cores to access the same instructions without needing to duplicate them across the core complex, as was the case with EPYC 2. “For these reasons, performance is improved across a broad range of applications that are just leveraging multi-core, multi-thread CPUs,” Beck said.

EPYC Security

Alongside the performance bump, EPYC 3 offers a bevy of new security capabilities built into its secure processor subsystem, which sits alongside the general-purpose processors on the die.

“Really the biggest change here, in Zen 3, was secure nested paging,” Clark said.

The capability is part of AMD’s secure encrypted virtualization-encrypted state (SEV-ES) feature set, which is designed for use in confidential computing environments where the memory needs to be encrypted in use. When enabled, secure nested paging helps to prevent a malicious hypervisor from executing replay or remapping attacks on the workload.

With this release, AMD also implemented stronger protections against Spectre-style side-channel attacks.

“When they hit on our first-gen EPYC, we were already in production, so we had to come up with solutions out of what really existed in our hardware,” Clark said.

With EPYC 2, which was launched in 2018, AMD was able to architect hardware mitigations for side-channel attacks, but Clark said it wasn’t until the third-generation that the company was able to implement high-performance mitigations.

Configurations and Performance

Like EPYC 2, the third-generation processor will be available in configurations ranging from 8-cores and 16 threads to 64 cores and 128 threads. New to this generation are 28-and 56-core configurations.

AMD has broken down each SKU into three broad categories that prioritize core performance, core density, or attempt to balance high core counts with high clock speeds.

In performance benchmarks provided by AMD, the chipmaker’s EPYC 3 processors went toe to toe with Intel’s highest-end, 28-core Xeon Scalable processors in SPECRate 2017. According to AMD, its highest-end 64-core processor offered more than twice the performance advantage over the Intel chip.

Peddibhotla touted the 28-core processor in particular as offering the same core density as Intel’s highest-end Xeon Scalable processors, while providing similar performance for less money.

However, it should be noted that AMD is comparing its third-generation chips against Intel’s Xeon Gold 6258, which is based on the old 14-nanometer Cascade Lake microarchitecture. Intel’s third-generation Xeon Scalable processors, based on the 10-nanometer Icelake microarchitecture, aren’t expected to reach the market until later this quarter. As a result, the performance deltas between the company's flagship chips are likely to diminish significantly once they reach the market.