Xilinx unveiled its most powerful data center accelerator yet at Supercomputing 2021 today. The chipmaker claims the FPGA can accelerate streaming data, high input/output (I/O) math, big data analytics, and artificial intelligence (AI) workloads while consuming less power and space than competing GPUs from Nvidia.
On paper, the Alveo U55C is just a more compact U280 with a bit more high-bandwidth memory (HBM). But according to Nathan Chang, high-performance compute (HPC) product marking manager at Xilinx, that’s kind of the point.
By doubling the amount of HBM to 16 gigabytes and reducing the accelerator’s footprint to a single-slot, half-length PCIe card, Chang claims Xilinx can achieve twice the compute density and four times the memory in the same footprint of the U280.
“When you’re building a node, you’re able to pack twice the compute into those two slots,” he said. The Alveo U55C “combines many features that today’s HPC workloads need. We’ve brought more HBM2 and key features from our previous cards all in a much slimmer, more efficient design.”
And this space savings allows the U55C to compete directly with Nvidia’s GPU-based accelerators, the company claims. Additional space savings are achieved by integrating networking functionality directly into the FPGA, which also has the benefit of reducing networking overheads on the host CPU.
A similarly equipped system using a GPU would require a separate NIC card consuming three PCIe slots to the U55C’s one, Chang explained.
Xilinx further reduced memory bottlenecks by boosting the HBM on each card, especially when working with large data sets. By running the entire workload in HBM, Xilinx is also able to dramatically reduce the number of system memory calls.
Xilinx Pushes Alveo to the MainstreamWith the launch of the Alveo U55C, Xilinx is looking to expand the use cases for FPGAs in the data center, where they’ll compete directly with Nvidia and AMD’s GPU-based accelerators.
“What GPUs did for the data center is they helped pave the way for more accelerators and ultimately more heterogeneous compute,” Chang said. “We’re shifting the position of the Alveo in the data center. It’s not just for niche architectures or specific data problems anymore.”
To achieve this, Xilinx spent the last several years abstracting the notoriously difficult process of developing for FPGAs. The result was Xilinx's Vitis software platform, which orchestrates the process of deploying and scaling workloads across multiple Alveo accelerators.
Vitis supports a wide range of AI frameworks and programming languages including Pytorch, Tensorflow, C, C++, and Python.
Xilinx claims another benefit of its FPGA-based accelerators is support for standard Ethernet networking using second-generation remote direct memory access over converged Ethernet (RoCE v2),
This has allowed Xilinx to bring forth an “Alveo network that competes with InfiniBand in performance and latency,” Chang said, adding that unlike InfiniBand, Alveo can be deployed in existing data center environments. “Anywhere there’s an Ethernet network, we’ll be able to plug these in.”
Early AdoptionAlso at Supercomputing 2021, Xilinx shared a handful of early HPC deployments of its new FPGAs.
One of the largest deployments of the cards is in Australia’s Commonwealth Scientific and Industrial Research Organization (CSIRO) Square Kilometer Array radio telescope.
The research group deployed 420 of Xilinx’s U55C FPGAs to accelerate signal processing for the 131,000 antennas that make up the massive telescope.
A key consideration behind the U55C’s adoption was compute density and power consumption, Chang explained. Because the satellite array is so remote, processing had to be done on location using a combination of solar and diesel generator power.
Ansys, the developer of a popular crash simulation software used by major automakers, has also extended support for the Alveo platform. Xilinx claims its U55C can achieve a fivefold improvement in performance compared to x86 CPUs alone in finite element method simulations.
Similarly, graph analytics software vendor TigerGraph is using multiple U55C cards to accelerate two of its most popular algorithms for recommendation and cluster engines. Xilinx claims that a cluster of Alveo U55C FPGAs achieved 45-times higher performance over a traditional CPU cluster.
The Alveo U55C is available now from Xilinx and as-a-service through public cloud providers.
Comments