The Nvidia Groq 3 LPU up close
An up-close Groq 3 LPU – Sebastian Moss

Nvidia’s LPX racks for AI inference accelerators have entered full production, the company confirmed.

Unveiled at GTC back in March, the rack-scale platform came about following Nvidia’s acqui-hire of eponymous startup Groq. LPX is liquid-cooled and houses 256 Groq 3 language processing units (LPUs) interconnected through 640 Tb/s of scale-up bandwidth.

Nvidia LPX LPU rack
A glimpse inside the LPX from GTC 2026 – Sebastian Moss

Inside the LPX rack itself are BlueField-4 data processing units (DPUs), Vera central processing unit (CPU) racks, and STX storage servers all tied together with the recently debuted Spectrum-6 Ethernet networking tech.

The platform is not a replacement for Nvidia’s flagship NVL72 platform, but rather a complementary add-on for operators wanting to power low-latency AI inference workloads.

During the Hot Chips event in Palo Alto this week, Nvidia cited industry benchmarking results that saw the LPX platform support 3,400 output tokens per second. The chip giant claims it can be used to drastically reduce the time it takes to perform agentic-related tasks from hours to mere minutes, offering four-times faster responsiveness for agents and latency-sensitive workloads.

Nvidia founder and CEO Jensen Huang said the LPX will “transform how intelligence is produced, delivering another giant leap in AI throughput, efficiency, and responsiveness.”

“Inference is the growth engine of AI. Nvidia Grace Blackwell and NVL72 revolutionized large language model (LLM) inference with an unprecedented leap in performance and efficiency,” Huang said. “Vera Rubin extends that vision with workload-optimized AI factory configurations designed for the era of agentic AI, advancing the performance frontier with LPX for ultrafast token generation.”

The LPX platform is expected to be available later this year.

Among its early adopters is neocloud darling Nebius, which plans to bring the Groq 3 LPX platform into its Token Factory offering. Nebius CTO Danila Shtan said the move will make “every step of an agent’s loop feel instant.”