AMD unveiled its Instinct MI100 GPU today in a bid to accelerate scientific research in high-performance computing (HPC) and artificial intelligence (AI) workloads. The chipmaker claims the GPU is capable of 11.5 teraFLOPs of performance in floating-point 64 tasks.
"Today, AMD takes a major step forward in the journey toward exascale computing as we unveil the AMD Instinct MI100," said Brad McCredie, VP of data center GPU and accelerated processing at AMD, in a statement.
Previously, AMD's presence in the data-center GPU market has been limited.
The MI100 is built using AMD's new Compute DNA (CDNA) architecture, which, when paired with the company's second-generation EPYC processors, can achieve a seven-fold performance increase, compared to previous generation AMD accelerators running floating-point 16 AI training workloads.
Each GPU boasts 120 compute units and 7680 stream processors. The GPU is fed by 32 gigabytes of high-bandwidth, error-correcting memory.
An Open Source FoundationAccording to McCredie, the MI100 squarely targets scientific research. To support this mission, AMD is making available a software platform designed to provide scientists the tools they need to optimize their HPC workloads.
The AMD ROCm software platform is an open source toolset consisting of compilers, programming APIs, and libraries. It is designed for use by exascale data scientists to develop applications for the MI100 GPU.
"The new chip is an attractive HPC platform and a respectable chip for AI training workloads, but the challenge any NVIDIA competitor must address lies in the software needed to engender an ecosystem of AI models and applications," wrote Karl Freund, senior analyst at Moor Insights and Strategy, in an article contributed to Forbes.
Freund called ROCm an elegant vision for enabling HPC application development, but notes that it will take a considerable amount of time for AMD to realize its goals.
However, early impressions appear positive.
"We've received early access to the MI100 accelerator, and the preliminary results are very encouraging. We've typically seen significant performance boosts up to [200% to 300%] compared to other GPUs," said Bronson Messer, director of science at the Oak Ridge Leadership Computing Facility, in a statement.
Messer expressed enthusiasm regarding AMD's decision to make the toolkit open source. "The fact that the ROCm open source platform and HIP develop tool are open source and work on a variety of platforms is something we have been absolutely almost obsessed with since we fielded the very first hybrid CPU/GPU system," he said.
The MI100 will begin shipping in Dell, Gigabyte, Hewlett Package Enterprise, and Supermicro systems by the end of the year.
CompetitionAMD's Instinct MI100 GPU faces stiff competition from rival Nvidia, which gave its own A100 data-center GPU a revamp today upping its video memory to 80 gigabytes.
According to Freund, the MI100 is AMD's first data center GPU to go toe to toe with Nvidia.
"Should NVIDIA be worried about AMD's new entry into the data center? In AI, no. In HPC, yes," he wrote. "AMD’s new GPU is an excellent stepping stone to its exascale platform. I expect AMD will be selected by price-sensitive supercomputer installations which may not have a tremendous need for bleeding-edge AI on the same platform in the short term."
In AI intensive workloads, Freund said Nvidia's A100 has a distinct advantage, especially in workloads that use quantization.
He added that Nvidia's latest A100 GPUs also offer more than twice the memory at substantially higher bandwidths. "This announcement widens the performance gap versus AMD, although at a higher price point," he wrote.
Comments