It’s not often that senior executives from a chip company celebrate when a rival unveils a new piece of hardware, but that’s the odd position Cerebras cofounder and CTO Sean Lie admitted finding himself doing after Nvidia unveiled the Groq 3 language processing unit (LPU) during its GTC 2026 event.
Watching the keynote unfold in the back of an Uber on his way into downtown San Jose, California, Lie said he and others in the car were “blown away” by the announcement, adding: “This is, like, the best thing ever for Cerebras.”
The reasoning for the CTO’s show of support? Lie viewed the announcement as validating a story the chip company has long told: that the real economic battleground in AI is ultra‑fast inference, where tokens served fastest command the highest value, and that static random-access memory (SRAM)‑rich, wafer‑scale architectures – essentially a single, dinner‑plate‑sized chip packed with memory so model weights sit right next to compute – are uniquely positioned to win that race.
“We've been trying to tell for years. And now we have Jensen [Huang] on stage telling the story for us,” Lie told a packed side event taking place alongside GTC.
While the Nvidia CEO effectively validated the giant chipmaker’s approach, Lie said he also may have inadvertently played into their hands.
“[Huang] essentially acknowledged that GPUs can’t really compete in that segment. Around the 400- to 600-tokens per second range – the high‑value, high‑speed part of the market – he basically said NVL72 runs out of steam; it just can’t play there … [because] there isn’t enough memory bandwidth. That’s exactly the story we’ve been telling for years," Lie said.
Much to Nvidia’s likely chagrin over how it positioned the $20 billion Groq licensing deal, Lie argued that this is precisely why his rival “bought” Groq – specifically to pair its GPUs with Groq’s hardware in its new LPX racks.
The Cerebras CTO again went on the offensive: “That can't actually make the Groq tokens any faster.”
“It can make the Groq tokens higher throughput, maybe lower the cost … the way Nvidia is using Groq is actually where the GPU is, like the accessory to the Groq chip, to be able to help with the throughput on the pre-fill side," Lie said.
Lie suggested that the Nvidia-plus-Groq performance would eventually max out somewhere between 500- and 1,000-tokens per second, while his own company’s flagship dinner-plate-sized hardware runs “1,000 tokens per second.”
“The reason is because of wafer-scale,” Lie said, contending that the next-generation Groq system revealed at GTC has 90-times less SRAM than Cerebras' current generation Wafer-Scale Engine 3 (WSE-3).
For a two-trillion parameter AI model, Lie estimated an enterprise would need access to a minimum of 2,000 Groq chips, compared to just 20 Cerebras chips.
“When you have thousands of chips, even if on paper you have memory bandwidth, the reason why this solution is only twice as fast as GPUs is because there's just so much overhead and interconnecting all together," Lie explained. "You've got to connect together thousands of chips, and every single one of those interconnects adds latency, adds overhead, and not to mention they are also interconnected to GPUs, which also adds latency.
“All of that integration overhead is solved on our single-wafer chip, and that's where it doesn't matter how much memory bandwidth you have; if you're now constrained by all of this other overhead, that becomes a problem, and that's the reason why our integration is so significantly faster than the Groq integration" Lie added.
Other Cerebras executives saw (at least half) of Jensen’s GTC two-hour-long keynote as a pitch for Cerebras’ worldview.
Cerebras CEO and fellow cofounder Andrew Feldman described the company’s hardware as being “so much faster, you had to change the y-axis in order to graph us,” arguing that what really matters in modern AI isn’t raw flops but how quickly you can move model weights from memory into compute – or the “size of the straw” on a cup.
Where a GPU's narrow, off‑chip memory pipes leave them fundamentally disadvantaged an the ultra‑fast inference race, Feldman explained Cerebras instead opted for wafer-scale to “pack fast memory onto a chip and have enough capacity to hold many models.”
“What determines how fast you can get Coke into your mouth? Only the size of the straw, that’s it," Feldman said. "And the size of the straw on GPUs is tiny, and it’s slow.”
Cerebras weren’t celebrating just the Groq announcement at GTC, as mere days before it received arguably its biggest industry validation when Amazon Web Services (AWS) signed it up to power an AI inference solution. That deal would see it partner with the hyperscaler’s Trainium-powered servers nd Elastic Fabric Adapter (EFA) networking to support its Bedrock platform in AWS data centers.
That deal came hot on the heels of a $10 billion deal with OpenAI to supply 750 megawatts (MW) of compute power. The ChatGPT firm had mulled acquiring Cerebras in a team-up with Tesla back in 2017, when Elon Musk was at the then non-profit firm.
For Feldman and the team, these deals, combined with Nvidia’s unintended show of support, were a sign that its architectural approach was the right one.
“We are now powering the No. 1 AI provider. We're now deployed in the largest AI cloud. And we achieved this by solving problems in a particular wafer-scale that the entire industry thought could never be solved,” Feldman added.
Comments