Cerebras logo on the front of its rack server solution
– Ben Wodecki/SDxCentral

Cerebras unveiled its own rack-scale platform powered by next-generation wafer-scale AI chips, claiming performance up to 30-times faster than graphics processing unit (GPU)-based rivals.

The move sees Cerebras shift away from a standalone appliance approach to a full-scale multiwafer rack architecture dubbed Nexus. The rack-scale platform is designed to house Cerebras' new WSE-3 Turbos and future generations of Cerebras chips, providing double the input/output (I/O) bandwidth to keep AI workloads running at high speeds. It also features a “wafer-scale backpack” mounted at the rear and attached vertically to the power array to provide direct liquid cooling around the chips.

Front panel view of Cerebras CS4
Front panel view of Cerebras' CS4 – Cerebras

The first offering is the CS-4, which features three WSE-3 Turbos – each packing four trillion transistors and 900,000 AI-optimized cores – and sees improved wafer-to-wafer communication compared to previous generations. Cerebras claims interconnection latency between Turbos is as low as just two milliseconds, which can support more than 1,000 tokens per second for AI models exceeding 10 trillion parameters.

On the networking side, Cerebras makes use of remote direct memory access (RDMA) over converged Ethernet version 2 (RoCE v2) to bridge its Wafer-Scale Engines. The standards-based protocol provides an open option to communicate directly with heterogeneous host systems and central processing unit (CPU) storage arrays.

The standards approach also opens the doors for users to employ alternative vendor hardware from the likes of Nvidia or AMD to act as a “prefill engine” – where the rival accelerator handles incoming prompts and readies the underlying model before handing it off to the Cerebras system to perform ultra-low-latency decoding and generate the response. That ties in with deals the chip firm has made with those players that allows operators to pair its Wafer-Scale Engines with systems inside platforms like Helios.

Wafer I/O interface in Cerebras Nexus
Cerebras' next-gen wafer I/O interface – Cerebras

CS-4 also makes use of Cerebras' “Direct Wafer Links” switch-free interconnect to link the new Turbo chips. Little technical information was divulged on the proprietary tech, though Cerebras said that the wafer chip’s I/O module can support open ecosystem connectivity.

At a system level, CS-4 provides 750 petaflops of sparse floating-point 16 (FP16) of AI compute and 7.2 Tb/s of I/O bandwidth. Engineers won’t have long to wait to get their hands on this speed machine either, with the first shipments of CS-4 set to commence this quarter.

“In AI, speed is productivity,” Cerebras CEO Andrew Feldman explained. “Historically, fast inference meant using smaller and less capable models. Cerebras CS-4 delivers industry-leading speeds on the largest frontier models, fundamentally changing the paradigm. Every aspect of the design has been optimized to deliver the highest speeds with massive throughput. With the CS-4, AI is so fast that it fundamentally reshapes product experiences.”