Nvidia has unveiled the Rubin CPX, a GPU specially designed for intense AI workloads that could create an entirely new class of networking challenges.
Revealed at the AI Infra Summit in California, the Rubin CPX – expected in late 2026 – is designed to handle what Nvidia described as “high-value AI use cases,” which require sizable context windows (the amount of text it can process in a single go), like advanced coding tasks with 100,000 lines or video processing and generation.
To handle intense applications, Nvidia claims a single Rubin CPX offers up to 30 petaflops of compute, while the rack-scale Vera Rubin NVL144 CPX platform delivers 7.5 times more AI performance than the current-gen GB300 NVL72 system.
Also notable is its attention capabilities, with the chip firm touting the server configuration as being three times faster than the GB300 NVL72 – meaning the Rubin CPX can boost an AI model’s ability to process longer context sequences without a drop in speed.
‘Purpose-built for massive-context AI’
Nvidia's annual release schedule sees Rubin CPX as a part of the Vera Rubin series, which is the follow-up to the Blackwell line.
The next generation of Nvidia chips uses a monolithic die design that packs NVFP4 computing resources to enable high performance and energy efficiency for AI inference tasks.
According to Nvidia CEO Jensen Huang, the CPX processors enable AI models to reason across millions of tokens simultaneously.
At this scale, offering 1.7 petabytes per second of memory bandwidth per rack, the hardware creates intense networking demands that exceed what traditional data center fabrics were designed to handle.
To meet those networking demands, the Rubin CPX integrates with Nvidia's ever-expanding – and profitable – interconnect ecosystem.
The Vera Rubin NVL144 CPX configuration, for example, supports either the InfiniBand-based Quantum‑X800 or Ethernet-based Spectrum-X platforms. It can also work in tandem with Nvidia's recently unveiled Spectrum-XGS technology, which enables operators to connect multiple data centers into unified 'AI super-factories' for distributed processing of massive workloads.
Nvidia claims the benefits could help on the cost front, suggesting the Vera Rubin NVL144 CPX could enable companies to monetize “at an unprecedented scale, with $5 billion in token revenue for every $100 million invested.”
“Just as RTX revolutionized graphics and physical AI, Rubin CPX is the first CUDA GPU purpose-built for massive-context AI, where models reason across millions of tokens of knowledge at once,” Huang added.
Emerging uses see Rubin CPX power video generation & AI agents
In addition to its networking functionality, the Rubin CPX line of hardware will be supported by Nvidia’s AI stack, including software solutions like Dynamo, which helps to scale AI inference by boosting throughput, as well as the Nemotron family of multimodal AI models.
Among the partners exploring how Rubin CPX can accelerate their applications were generative AI platform Runway, which plans to use the hardware to power its video generation systems, and Magic, a firm developing foundation models to power AI agents for automated software engineering.
“With a 100-million-token context window, our models can see a codebase, years of interaction history, documentation, and libraries in context without fine-tuning,” said Eric Steinberger, CEO of Magic. “This enables users to coach the agent at test time through conversation and access to their environments, bringing us closer to autonomous agentic experiences.”
“We see Rubin CPX as a major leap in performance, supporting these demanding workloads to build more general, intelligent creative tools,” said Cristóbal Valenzuela, CEO of Runway. “This means creators, from independent artists to major studios, can gain unprecedented speed, realism, and control in their work.”
Comments