Generative Artificial intelligence (AI) is now the center of attention for performance engineering.


Today, MLCommons® announced new results for its industry-standard MLPerf® Inference v5.0 benchmark suite, which provides machine learning (ML) system performance benchmarking in an architecture-neutral, representative, and reproducible manner. The results indicate that the AI community is concentrating its efforts on generative AI scenarios, with hardware and software advancements leading to performance improvements over the past year.

The MLPerf Inference benchmark suite measures the speed at which systems can run AI and ML models across various workloads. This open-source and peer-reviewed suite creates a competitive environment that drives innovation, performance, and energy efficiency industry-wide. The latest MLPerf Inference results include new tests for Llama 3.1 405B, Llama 2 70B Interactive, RGAT, and Automotive PointPainting for 3D object detection.

Generative AI scenarios have gained traction, as submissions for the Llama 2 70B benchmark test have increased 2.5 times in the last year. The median submitted score for Llama 2 70B has doubled compared to the previous version, and the top score is 3.3 times faster.

David Kanter, head of MLPerf at MLCommons, stated, “It’s clear now that much of the ecosystem is focused squarely on deploying generative AI, and that the performance benchmarking feedback loop is working. The community is setting new records for generative AI inference performance.”

This set of benchmark results includes submissions from six new processors: AMD Instinct MI325X, Intel Xeon 6980P “Granite Rapids,” Google TPU Trillium (TPU v6e), NVIDIA B200, NVIDIA Jetson AGX Thor 128, and NVIDIA GB200.

Also new in Inference v5.0 is a benchmark utilizing the Llama 3.1 405B model, consisting of 405 billion parameters and supporting input and output lengths of up to 128,000 tokens. This benchmark tests three tasks: general question-answering, math, and code generation.

Miro Hodak, co-chair of the MLPerf Inference working group, remarked that the new benchmark reflects a trend toward larger models, which can enhance accuracy and versatility. The release also introduces an interactive version of the Llama 2 70B test, emphasizing response metrics crucial for query systems and chatbots.

A new datacenter benchmark that employs a graph neural network (GNN) model has also been introduced, useful for applications like recommendation systems and fraud detection.

The Inference v5.0 suite reports 17,457 performance results from 23 submitting organizations, including newcomers CoreWeave, FlexAI, GATEOverflow, Lambda, and MangoBoost.

According to Kanter, the MLCommons community is expanding capabilities in machine learning and is poised to deliver valuable performance data reflecting the rapid advancements in the field.

To view the results for MLPerf Inference v5.0, visit the Datacenter and Edge benchmark results pages.