Chip market analyst firm SemiAnalysis in its latest AI inference benchmark testing called out AMD, claiming the Nvidia rival’s inference optimizations were “not as competitive as one would expect” when working in tandem.
The second edition of the research firm’s InferenceMAX benchmark, now known as InferenceX, saw Nvidia’s GB300 NVL72 system as the king of the hill, boasting performance levels of up to 100-time over a strong H100 baseline in some configurations on what was an apparent first third-party evaluation of the rack-scale system.
But for AMD, SemiAnalysis’s InferenceX data told a different story. The mammoth report found AMD's individual inference optimizations performed well on their own, but combining them within its software stack produced results that fell well short of expectations.
But it was a different story when looking at what SemiAnalysis likened to "composability." Its researchers found that inference on AMD’s systems in line with its software across three major optimizations – disaggregated prefill, wide expert parallelism (wideEP), and four-bit floating point (FP4) – was where Nvidia looks to be pulling ahead.
Further compounding AMD's position was the fact that performance results showed single-node performance of its MI355X on FP4 was decent, but compared to Nvidia’s B200, a GPU released back in 2024, it was outclassed on several other optimizations. The analyst firm went as far as to channel their inner "Generation Alpha," claiming the AMD chip “gets absolutely mogged” by the rival system.
Brad Shimmin, VP and practice lead for data intelligence, analytics, and infrastructure at Futurum, said the composability issue “speaks to a lack of optimization between the selected model, model architecture, inference libraries, and the underlying chip rather than any shortcoming in the chip itself.”
“You can lease a Ferrari, but if you load it up with your friends for a hot lap, don't be surprised if your times fall a bit short of what's possible,” the analyst added.
Despite the admonishment, AMD’s performance results are not unfixable, with SemiAnalysis confirming that upon raising its concerns with the vendor, AMD plans to focus on software composability of FP4+distributed inferencing “across their whole software stack” – though such efforts won’t occur until after the Chinese New Year, as the majority of its related engineers are based in China.
The software problem
AMD has played semiconductor second fiddle to Nvidia’s crown, leaning on its ROCm software stack as an open source enticement for operators looking to avoid being locked into the market leader’s ecosystem.
The firm has made significant strides to try and close the gap, with a complete design overhaul to its Instinct line culminating in the MI355X, which the chip designer claims provides a 50% performance increase compared to its MI300Xs.
But on the last SemiAnalysis list, AMD’s MI355Xs were outpaced by Nvidia’s GB200 NVL72s across metrics like throughput-per-dollar and tokens-per-megawatt.
“The composability of disaggregated prefill, wideEP, and FP4 inference optimizations needs significant improvement,” the research team wrote. “While performance is competitive on AMD when enabling just a subset of the state-of-the-art inference optimizations, enabling all three major optimizations that labs use, AMD’s performance is currently not competitive with Nvidia’s.”
Brendan Burke, research director for semiconductors, supply chain, and emerging tech at Futurum, commented: "AMD's GPUs may not excel on all inference workloads yet remain useful to eight of the top 10 AI labs who know when to use them as part of a portfolio approach to AI compute.
"AMD shines in a single node configuration when high-bandwidth memory can store all the weights needed to run a workload, whether in training or inference" Burke continued. "Customers use this configuration with rapidly improving ROCm software to contain GPU costs and avoid system performance risks of disaggregation."
Blackwell Ultra’s big debut
While AMD found itself the subject of some criticism, for the folks over at Nvidia, it was just another day, with data showing software optimizations combined with its latest generation Blackwell Ultra line resulting in some serious performance capabilities.
For a first evaluation outing, the Blackwell Ultra-based GB300 saw some impressive results, chiefly its cost performance levels, in that it dramatically lowered cost per-million tokens at low latency compared with Hopper generation platforms.
This performance gain translates into superior economics, with Nvidia GB300 lowering costs compared with the Hopper platform across the entire latency spectrum. The most dramatic reductions occurred at low latency regimes, notably where agentic applications operate, with cost per-million tokens dropping by as much as 35-times versus Hopper. At higher interactivity levels and smaller batch sizes, however, that advantage narrows as the scale-up benefits of NVL72 are less pronounced.
But it wasn’t all Nvidia’s way this time out. SemiAnalysis found that for low batch size large scale-up workloads, the NVL72 iteration had a similar performance level as the 8-GPU module B300s. Despite a sizable bandwidth advantage for the larger iteration, SemiAnalysis found it doesn’t matter, because not even the much slower InfiniBand link is “saturated by the tiny batches of tokens in flight.”
On NVL72 losing its scale-out advantage versus 8-GPU modules, Burke said, "We are seeing that conventional 8-GPU HGX clusters remain highly useful for inference, as not all AI apps use high batch sizes in an effort to control costs. NVL72 racks have been initially used by AI labs for large-scale training runs and may also be used for scaled inference for high-value products.
"With Rubin around the corner, Blackwell racks will transition from training to inference and require software optimization to ensure performance gains across batch sizes," Burke added. "The industry must continue to prove customer willingness to pay for large batch inference to justify the long-term ROI."
Even more tests to come
With the second iteration of SemiAnalysis’s benchmarking behind us, the analyst firm pledged to further build out testing.
Future editions are expected to feature results covering hyperscale custom silicon, including Ironwood TPUs from Google and Amazon’s Trainium3s.
The firm also plans to add performance testing across more popular Chinese frontier models, including the highly anticipated DeepSeekv4, with industry rumors suggesting it’ll drop sometime this spring.
The Chinese lab that shook the tech industry last year has already debuted one offering in 2026: Engram, an architectural concept that essentially bypasses GPU memory constraints to allow firms running large language models (LLMs) to scale parameters aggressively without hitting GPU memory walls.
Comments