Ampere Computing’s acquisition of OnSpecta last week was a warning shot to Intel and Nvidia as the Arm-CPU vendor turns its attention to artificial intelligence (AI) inference.
While neither company disclosed the terms of the deal, the acquisition buys Ampere exclusive access to OnSpecta’s deep learning software (DLS) acceleration platform and AI modeling libraries.
OnSpecta’s DLS accelerates AI workloads by ensuring that the optimal processor instruction sets are called for a given AI framework. In addition to broad hardware support, DLS also boasts support for the most popular AI frameworks including TensorFlow, PyTorch, ONNX, and BERT.
For Ampere, DLS ensures that the Arm instruction set used by its Altra and Altra Max CPUs are utilized to their full potential, explained Ampere CPO Jeff Wittich, in an interview with SDxCentral.
“We were seeing a 4X-plus performance improvement, both in throughput and latency,” he said, adding that because DLS functions as an acceleration layer running between the AI framework and the hardware, it doesn’t require additional action by the user to implement, “they just get a big enhancement in performance.”
Ampere leaped into the spotlight early last year with the launch of its Ampere Altra CPUs, a family of Arm-based processors designed to go toe-to-toe with Intel and AMD in the data center and cloud arenas.
Since their announcement, Altra has seen broad adoption across the industry with Oracle Cloud, Microsoft Azure, and Cloudflare. Oracle was also among the first to pair OnSpecta's DLS platform with Ampere's CPUs earlier this year.
Ampere Amps AI InferenceAmpere claims the combination of DLS and its Altra CPUs allows it to compete directly with the best Nvidia and Intel have to offer, at least when it comes to AI inferencing.
While AI training is a much “sexier” market segment and gets a lot of attention, inferencing, particularly at the edge, is a far larger opportunity, Wittich explained.
And this is where OnSpecta’s software libraries come into play. “It allows us to exceed the performance of commonly used GPUs for inference, like the [Nvidia] T4,” Wittich said. “With a fraction of the cores that we have on a single Ampere Altra processor, we can beat the performance of a single T4.”
Scaling this up, he explained that a single Altra Max CPU with 128 cores could exceed the performance of four Nvidia T4 GPUs in AI inference workloads while consuming a fraction of the power.
From an economic standpoint this makes Altra extremely attractive compared to GPU-based inference, he added. “The idea of putting GPUs in all of your servers on the off chance that inferencing is done some part of the day is kind of a poor economic decision.”
Since the launch of Intel’s second-generation Xeon Scalable family, the company has touted its integrated AI inference capabilities on a loop.
OnSpecta “allows us to reach parity and exceed the performance of Intel,” without relying on specialized compute blocks for big vector instructions, Wittich said in a direct shot at Intel’s reliance on the AVX 512 instruction set. We “use our transistor area for stuff that can be used across all workloads, versus adding in specialized blocks that are inherently inefficient because they can't be used for everything.”
Only the Beginning of Ampere’s AI AmbitionThe OnSpecta acquisition is only the start of Ampere’s AI ambitions. The company is already working with the OnSpecta team to develop a library of pre-trained AI inferencing models, which it calls Model Zoo, for customers that are starting from scratch or don’t have an existing dataset to train.
“We’ll have dozens and dozens of models that have already been pre-trained for specific types of activities, and that way if somebody wanted to kickstart things they could just grab a pre-trained model and then get going,” he said.
Looking beyond software, Ampere is also building AI optimizations into its upcoming core design, which the chipmaker detailed earlier this year.
“As we look forward with our own core … we have a lot of opportunities to go in and optimize things at the micro-architectural level,” Wittich said.
Comments