A startup led by Lambda’s former COO that’s looking to build more efficient AI inference hardware raised $875 million in a Series C round.
Positron AI secured the cash across two tranches – a $375 million raise co-led by a host of investors including Atreides Management and Valor Equity Partners, and a $500 million Series C-1 round led by New Enterprise Associates and Netscape co-founder Jim Clark. Qatar's sovereign wealth fund, Cisco Investments, and Boardman Bay Capital Management were also among the backers, with the startup valuing itself at around $5 billion.
Positron was founded in April 2023 by neocloud alums Mitesh Agrawal, Thomas Sohmers, and Edward Kmett. It's designing what it describes as “memory-first inference systems” designed to skirt memory bandwidth and power constraints.
The startup claims its custom silicon line, dubbed Asimov, “realize[s] more than 90% of their available memory bandwidth,” compared to what it claims is just under 30% for graphics processing units (GPUs) running the same models.
The team opted for low-power double data rate 5X (LPDDR5X), a low-power dynamic random access memory (DRAM) primarily designed for mobile and edge devices not currently suffering the same constraints befalling high-bandwidth memory (HBM). Per-chip memory capacity ranges from 277 gigabytes to a staggering 2.3 terabytes, with Asimov capable of being interconnected into clusters containing up to 16,384 chips.
Asimovs are housed in Positron’s Titan inference server platform, which is designed for both air- and liquid-cooled deployments. The server features either four or eight of the custom chips, with the startup claiming it can support up to 32 trillion parameters per server. Chip-to-chip bandwidth equates to around 128 Tb/s, providing the massive memory bandwidth needed to keep packets flowing when running high-end models.
"The Positron inference architecture balances compute, the enormous required memory bandwidth, and extraordinarily large context and weight storage. Its optimized power consumption, cost, and density is near ideal for inference requirements of the next evolution of frontier large language models (LLMs) with trillions of parameters," Netscape co-founder Clark said.
The founding team previously helped upend the cloud infrastructure market and is now targeting the conventional chip industry.
Agrawal spent more than seven years at neocloud darling Lambda, leading AI and machine learning GPU cloud revenue and operations. Among other entries on his CV was a short stint at VMware as a software consultant.
Kmett also spent time at Lambda as a hardware architect before joining Groq – the language-processing unit (LPU) developer ravaged by an Nvidia acqui-hire. Sohmers too was previously at Groq, notably serving as its head of technology and architecture.
The $875 million raised will go toward Asimov tapeout, production ramp of its Titan servers, and securing LPDDR5X capacity. Asimov is set for tapeout on Taiwan Semiconductor Manufacturing Co.’s (TSMC) advanced three-nanometer (3nm) process at the end of 2026, with production in the second half of 2027.
“Speed matters in this market, both in how quickly we ship new generations of silicon and in how quickly they reach customers," Positron CEO Agrawal noted. "Deploying Atlas at scale taught us an enormous amount about what inference customers actually need, and we have carried those lessons directly into Asimov and Titan. Our focus now is to tape out Asimov, bring Titan to production, and scale manufacturing to meet the demand in front of us. This financing gives us the resources to do exactly that.”
Comments