Rivals Intel and Nvidia both touted support for the latest large language model (LLM) from Meta, signaling the competitive pressures of the high performance computing (HPC) infrastructure market.
Effective on the day Meta Llama 3 launched, the GPU market underdog claimed validated support for Llama 3 8B and 70B models across its artificial intelligence (AI) product portfolio, which includes Intel Gaudi accelerators and Intel Xeon processors.
According to Intel VP and GM of AI Software Engineering Wei Li, “Meta Llama 3 represents the next big iteration in large language models for AI,” and Intel is keen to help enterprises use Llama 3 to “develop products for cutting-edge AI applications.”
Intel tested the performance of the Llama 3 8B and 70B models on its hardware using open source software like PyTorch and DeepSpeed. The tests found the company’s Xeon 6 processors with Performance-cores (P-cores) improve inference latency by 2x when compared to its fourth-generation Xeon processors. The tests also demonstrated Xeon 6 can run Llama 3 70B at speeds of less than 100 milliseconds per generated token.
Nvidia's take on Llama 3Nvidia also announced support for the new LLMs in its TensorRT-LLM open-source software library, which improves inference performance on Nvidia GPUs.
Nvidia is “a good competitor,” Intel CEO Pat Gelsinger said following the closely timed announcements of Intel’s Gaudi 3 GPU and Nvidia latest Blackwell GPUs.
While Nvidia’s GPU sales represent more than 95% of global 2023 GPU revenue, Intel’s speed to support new ecosystem developments highlights the company’s fervor to catch up, which analysts concur is plausible.
Intel “could gain some share from Nvidia this year,” Dell’Oro Sr. Research Director Baron Fung told SDxCentral.
In the latest round of MLPerf testing, Nvidia submitted TensorRT-LLM and found the software nearly triples LLM inference performance on Nvidia GPUs. TensorRT-LLM was also key to the vendor’s high performance on the Llama 2 70B test.
Nvidia TensorRT-LLM support for Meta Llama 3 is available through a browser user interface or through API endpoints running on a stack of software from the Nvidia API catalog. Llama 3 is packaged as a Nvidia NIM within the vendor’s AI platform, which comes with a standard API for deployment in local workstations, public cloud or on-premises data center environments.
Comments