Chip startup Habana Labs announced its second product: an artificial intelligence (AI) processor named Gaudi that is purpose built for machine learning (ML) training.
Gaudi comes just months after Habana started shipping its first AI processor. The first product targets inference, and if you guessed it was also named after a famous dead Spanish artist, you’d be right. The startup’s original AI chip is called Goya, and Facebook uses this inference processor for its Glow ML compiler.
These two chips comprise Habana’s product strategy — and its plan to steal market share from AI chip giant Nvidia.
“The market has been dominated by Nvidia and customers view Nvidia as really having a lock on them. Frankly, our customers are hungry for an alternative even if [our processors] weren’t so much better,” said Eitan Medina, chief business officer at Habana. Benchmark tests show both Goya and Gaudi perform better and faster than Nvidia GPUs and Intel CPUs, he added. “Our motto is AI performance, not stories. We are talking about actual application performance, and we are giving you hardware with software that you can test.”
There are two components of AI: training and inference. The training phase, as the name suggests, trains the ML system to better understand a data set. Inference is when a machine acts on a new data sample to infer an answer to a query.
Gaudi ProcessorsAccording to Habana, training systems based on Gaudi processors will deliver an increase in throughput of up to four times over systems built with equivalent number GPUs. And Gaudi maintains high throughput even at smaller batch sizes.
This is important because training is often done at scale, “and as you scale you need each processor to maintain throughput even as you give it a smaller and smaller batch,” Medina said. So this high throughput allows performance scaling of Gaudi-based systems from a single-device to large systems built with hundreds of Gaudi processors, he explained.
Gaudi also has on-chip integration of RDMA over Converged Ethernet (RoCE v2) functionality within the AI processor. This enable the scaling of AI systems to any size using standard Ethernet. This is also a shot at Nvidia because it means Habana customers can use standard Ethernet switching — as opposed to Nvidia GPU-based systems, which rely on proprietary system interfaces — for both scaling-up and scaling-out AI training systems. Additionally, the Gaudi AI training processor supports the Open Compute Project (OCP) Accelerator Module (OAM) specification.
“With Gaudi, a data center avoids locking themselves in to any processor,” Medina said. “Let the processor guys compete on processors and let’s make sure all interfaces and form factors are standard.”
The Gaudi processor includes 32 gigabytes of HBM-2 memory and is offered in two forms:
- HL-200 – a PCIe card supporting eight ports of 100Gb Ethernet.
- HL-205 – a mezzanine card compliant with the OCP-OAM specification, supporting 10 ports of 100Gb Ethernet or 20 ports of 50Gb Ethernet.
Habana is also introducing an 8-Gaudi system called HLS-1, which includes eight HL-205 Mezzanine cards, with PCIe connectors for external host connectivity and 24 100Gbps Ethernet ports for connecting to off-the-shelf Ethernet switches, which allows scaling-up in a standard 19-inch rack by populating multiple HLS-1 systems.
Can It Unseat Nvidia?“I’ve been following Habana closely since last September when they announced their first product, the Goya inference chip,” said Karl Freund, an analyst at Moor Insights and Strategy. “The company seems to have the ability to, No. 1, produce a really attractive design, and No. 2, to be able to execute on that design. But it takes more than a good chip to unseat Nvidia in this space.”
Nvidia started shipping its Volta AI chip two years ago, and the vendor will likely announce a Volta’s successor later this year, Freund said. Plus Intel and Facebook are working on Nervana, a processor for inference, and Qualcomm and Arm have announced plans to develop AI chips as well. There are also at least a dozen startups pursing the same market.
“When you look at the space that’s this crowded you have to ask yourself how many companies can the market sustain? With companies like Intel and Nvidia and Qualcomm, I think the answer is: at most, only a handful will be successful,” Freund said. “Habana will need to have not only a good chip but a good roadmap. It’s going to be a lot of fun to watch.”
Comments