Nvidia unveiled it’s H100 GPU and Hopper compute architecture, the successor to its wildly popular Ampere-based A100 artificial intelligence (AI) accelerator during Tuesday’s virtual GTC event.

The new chip is massive, packing 80 billion transistors, which are based on a custom 4-nanometer manufacturing process developed in collaboration with Taiwan Semiconductor Manufacturing Co. (TSMC). It features Nvidia’s fourth-generation Tensor Core architecture and, according to Nvidia, uses a combination of TSMC’s chip-on-wafer-on-substrate packaging technology and third-generation high-bandwidth memory (HBM3) to achieve substantial performance gains over its predecessor.

“The Hopper architecture is a giant leap over Ampere,” CEO Jensen Huang said, adding that H100 is six times more performant than its predecessor in AI workloads. It’s “our largest generational leap ever.”

This is achieved thanks to a combination of hardware and software advancements, including an improved Tensor Core architecture and support for 8-bit floating point calculations. The Hopper architecture also introduces a new instruction set designed to accelerate dynamic programming algorithms.

“Many real-world algorithms grow with combinatorial or exponential complexity, examples include the famous traveling salesperson optimization problems,” Huang said, explaining that the technology breaks down complex problems into smaller pieces that are then solved recursively, reducing time required by 40-times over.

The result is models that used to take weeks can now be completed in a matter of days, he added.

These performance gains aren’t without compromise, however. Nvidia Hopper architecture is its hottest and most power hungry to date.“Designed for air and liquid cooling, the H100 is also the first GPU to scale in performance to 700 watts,” Huang said.

Nvidia Brings Confidential Computing to the GPU

Beyond raw compute potential, Nvidia’s Hopper architecture is its first with native support for confidential computing.

“Confidential computing today is only CPU based,” Haung said. “Hopper introduces the first GPU confidential computing.

Like the A100 before it, each H100 can be partitioned into seven instances using Nvidia’s multi-instance GPU (MIG) technology announced in 2020. With Hopper, customers gain per-instance isolation and input/output virtualization. What's more, each instance can support a different tenant, something that wasn't possible in previous generations.

“Each Hopper multi-instance supports confidential computing with trusted execution environment,” he said. “Software developers and services can now distribute and deploy their proprietary and valuable AI models on shared or remote infrastructure, protecting their intellectual property and scaling their business models.”

News DGXs Arrive Alongside EOS Supercomputer

Of course, with new GPUs comes new DGX servers. And this year’s release introduces a bevy of new features designed to eliminate performance bottlenecks in large compute clusters.

Nvidia’s DGX H100 shares a lot in common with the previous generation. It features eight H100 GPUs connected by four NVLink switch chips onto an HGX system board. Each DGX features a pair of “Gen. 5 CPUs” alongside eight ConnectX-7 smartNICs for a total of 3.2 Tb/s of aggregate networking per system.

If more performance is required, Nvidia's new NVLink Switch System enables customers to interface up to 32 DGX systems into single 256 GPU cluster. And for customers looking to deploy even larger clusters, 140 of these systems can be deployed as part Nvidia’s DGX Superpod using the company's previously announced Quantum-2 InfiniBand switch.

In fact, the DGX Superpod is the blueprint on which Nvidia’s latest supercomputing project dubbed EOS will be built. When it comes online later this year, EOS will deliver 275 petaFLOPS of conventional compute performance and 18.4 exaFLOPS — four times that of the Fugaku supercomputer — in AI workloads, the company claims.

“We expect EOS to be the fastest AI supercomputer in the world,” Huang said.

H100 Gains Onboard Networking

For deployment in conventional servers, Nvidia also offers the H100 in a standard PCIe form factor with an integrated ConnectX-7 smartNIC.

According to Huang, onboard networking is necessary to mitigate system-level bottlenecks. “Moving data to keep the lightning fast GPUs fed is a serious concern,” he said. “Moving data in traditional servers overloads CPU and system memory and are bottlenecked by PCIe.

If this sounds familiar, Nvidia previously offered a version of its BlueField-2 data processing units (DPUs) with an integrated A100 GPU.