Graphic depicting Nvidia versus AMD
– Adobe/Edits by SDxCentral

While the eyes of the tech world were firmly affixed on Nvidia last week for its GTC event and the unveiling of its new Groq language processing unit (LPU), its big rival doesn’t look to be sitting back as AMD went on the offensive.

In the weeks leading up to the showcase that some proclaim to be the "Super Bowl of AI," a SemiAnalysis report had Nvidia as the top dog in the AI inference software space, calling out AMD for a "composability" gap.

Nvidia, of course, didn't hold back, as CEO Jensen Huang styled himself the "inference king" holding a JPEG of a wrestling belt adorned with the title – a moniker the company then again showboated in a strange AI-generated song that sealed his GTC keynote.

AMD didn’t let the findings slide, however. In a press roundtable shortly before GTC, Anush Elangovan, AMD’s VP for AI software, said updates have since seen performance of AMD’s software stack having “almost caught up” or even outpacing Nvidia’s B200 on four-bit floating point (FP4).

While Elangovan said AMD took on board some of the analyst firm’s feedback around FP4 performance, the reality for his team was that “most” of its customers serve in eight-bit floating point (FP8).

“We focus on where our customers are,” Elangovan said in response to an SDxCentral question on the findings. “Another way to validate that is if you go to OpenRouter and sort by the model, you will see that there's like one FP4 model that is at the top of the list. Everything else is FP8. So FB8 is still king of the hill, and definitely there'll be a transition to FP8. It's still not the case there.”

Elangovan’s defense of AMD’s Radeon Open Compute (ROCm) software stack coincides with its rival marking the 20th anniversary of its Compute Unified Device Architecture (CUDA) platform.

AMD has long sought to position ROCm, which only released in 2016, as the open alternative to Nvidia’s stack, giving developers and enterprises a means to run AI and high-performance computing (HPC) workloads with popular frameworks like PyTorch, TensorFlow, and JAX upstream, while also moving models away from Nvidia without having to rewrite the underlying code.

While Nvidia came out late last year with CUDA Tile as an effort to essentially democratize writing and making programs for GPUs across large datasets, Elangovan said AMD was “heavily” invested in Triton – an open, Python-first GPU compiler originally developed at OpenAI – viewing CUDA Tile and related efforts as “a direct response to Triton’s increasing democratization of programming.”

AMD, Elangovan argues, wants Triton to become “the de facto” high-level abstraction so that moving from Nvidia to AMD is “zero friction,” with AMD’s own lower-level domain-specific languages (DSLs) – such as Fly DSL and Wave – handling more hardware-specific tuning under the hood.

Triton users can write kernels as decorated Python functions, while the tool compiles them down to efficient GPU code. It gives AMD users the ergonomics of high-level Python without their GPUs being locked into a single hardware stack.

And just like CUDA having years on ROCm, Triton debuted in 2021 – with four whole years of industry developers having gotten used to it prior to CUDA Tile launching.

“We are pushing very hard for Triton to be the de facto, because if Triton gets adopted, the abstraction level is as high as possible. And it's ... zero friction to move from Nvidia to AMD,” Elangovan added.

Defending the CPU moat from Nvidia’s encroachment

AMD must have caught wind that Nvidia would make its Vera central processing unit (CPU) one of the headline showcases at GTC, as AMD published a blog championing its own processor line mere days before the GTC event kicked off.

Nvidia’s push into the CPU space with its Grace platform was only ever going to draw competitive ire from AMD, with the latter having short work of usurping Intel’s once-dominant foothold in the market.

Perhaps evidence of that ire, AMD’s aptly timed blog post cited SPEC CPU Benchmark data that its 5th-Gen EPYC CPU offered 2.1-times higher performance per core against Nvidia’s Grace Superchip systems, while up to a 2.26-times uplift in operations per watt. And lest we forget the top two supercomputers in the world – El Capitan and Frontier – leverage AMD processors.

AMD 5th Gen Epyc
AMD's 5th Gen EPYC CPU – AMD

Before Huang and Nvidia could get a chance to pitch its CPU line as a means to power the emerging glut of agentic AI workloads, AMD beat them to it.

“Think of the relationship between CPUs and GPUs in AI data centers as that of a head coach with a team of agile athletes,” the AMD post reads. “The CPU head coach calls the plays, reacts to the other team, watches the clock, and keeps all the players moving in the right direction. GPUs are the players, each of them specializing in one very efficient part of one play at a time.“

AMD is looking to position its server CPUs as a means to orchestrate GPUs and handle complex agentic tasks ready for its larger hardware brethren. The AMD post goes as far to say: “GPUs with their smaller cores are designed for simpler chores that they perform again and again at a rapid pace.”

Hyperscalers have already taken notice as they look to ready theemselves against the agentic onslaught.

Microsoft lined up AMD’s Turin line of processors to further broaden the underlying hardware available for its virtual machines (VMs). The Da/Ea/Fasv7-series VM families, launched in early February, feature CPUs that the pair claim offer higher instructions-per-clock, greater memory bandwidth, and support for advanced vector instructions, on top of 35% better CPU performance compared to the prior v6 AMD-based generation.

That follows Google Cloud’s earlier adoption of AMD’s 5th-Gen EPYC processors for its C4D and H4D VM instances powering AI inference, HPC workloads, and general-purpose computing workloads.

While Nvidia finds its feet in the CPU space, AMD is already making strides to secure the edge. The vendor launched edge-centric subsets of its enterprise-grade CPUs last September, touting them as a means to power latency-critical applications.

Nvidia is busy flexing its software muscles, but looks to be facing stiff competition as it wades into AMD’s entrenched and well-defended processor market.