Less than a year after Nvidia CEO Jensen Huang hefted 50 pounds of GPU from his oven, the chipmaker is back with an even more capable chip designed for larger and more complex artificial intelligence (AI) and machine learning models.

The refreshed A100 GPU announced today boasts twice the memory of its predecessor, and it offers 80 gigabytes of high-bandwidth memory capable of feeding the GPU cores at 16 Tb/s.

This might sound like overkill, but according to Nvidia, large pools of video memory are essential in AI training workloads, like recommender models. These models which often have massive tables representing billions of users and billions of products.

"We are announcing the A100 80 gigabyte GPU to continue to advance our AI supercomputing platform and enable researchers to tackle the world's most important scientific and big data challenges," said Paresh Kharya, senior director of product management and marketing at NVIDIA.

By doubling the available video memory, Nvidia claims it can run larger workloads across a single DGX A100 server. Or, on the flip side, the larger pool means more memory per virtualized GPU instances. By taking advantage of Nvidia's multi-instance GPU (MIG) technology, each A100 can now be segmented into 7 GPUs, each with 10 gigabytes of video memory.

"If you look at data analytics, you just cut the price of data half with this system," said Charlie, VP and GM of DGX systems at Nvidia.

The 80-gigabyte versions of the A100 GPU will be available in Nvidia's DGX A100 server beginning this quarter.

According to Nvidia, partners including Atos, Dell Technologies, Fujitsu, Gigabyte, Hewlett Packard Enterprise, Inspur, Lenovo, Quanta, and Supermicro will begin offering systems built around the updated chip beginning in the first half of 2021.

An AI Supercomputer Under Your Desk

Launching alongside the upgraded GPU is the DGX Station A100. The workstation packs four A100 GPUs of either the 40 gigabyte or 80-gigabyte varieties into a desktop-style form factor.

"This is a supercomputer under your desk. The same HGX board, that [Kharya] talked about being in the world's fastest supercomputers, is in this system," said Boyle. "Through an incredible feat of engineering, our mechanical designers, electrical designers, thermal designers were able to come together and build a workgroup appliance for multiple users that plugs into any wall socket around the world."

Nvidia claims the workstation is the only petascale workgroup server of its kind and is capable of delivering 2.5 petaFLOPs of AI performance. Additionally, by taking advantage of Nvidia's MIG technology, each DGX Station can run up to 28 separate workloads spread across its four A100 GPUs.

According to Boyle, the new DGX Station will enable customers to begin working on AI models while their supercomputing cluster is under construction.

"Typically, when customers want to build a supercomputer, they plan and then it takes months or years to build. And the first thing their users ask is when can I get the first node so I can do software development," Boyle said. The DGX station can serve as that first node because it is the same underlying infrastructure used by all A100-based systems, he explained.

In fact, the system is already in use by several large enterprises including BMW Group Production, the German Research Center of AI, Lockheed Martin, NTT Docomo, and the Pacific Northwest National Laboratory.

Nvidia Lofts InifiBand to the Exascale

The A100 isn't the only hardware getting a spec bump. Nvidia today upgraded its Mellanox InfiniBand network interconnects from 200 Gb/s to 400 Gb/s and introduced "in-network computing engines" designed to provide additional acceleration for workloads spread across multiple GPUs.

"The need for bandwidth is there," said Gilad Shainer, SVP of networking at Nvidia, adding that to feed the A100 GPUs and consequently the AI models 400 Gb/s networking is necessary.

Shainer added that the updated InifiBand network interconnects offer more than higher bandwidth. Their integrated compute engines analyze data on the fly to optimize workloads spread across large clusters. According to Nvidia, offloading operations to its InfiniBand NICs can accelerate deep learning and training operations by as much as 3,200%

While faster fiberoptic interconnect technologies now exist within the datacenter, Nvidia is achieving 400 Gb/s speed using copper connections.

The new InifiBand NICs are expected to launch during the second quarter of 2021.