Meta
– Getty Images

Meta became the latest advocate of memory recycling in efforts to stave off the AI-driven "RAMpocalypse."

As reported by The Register, the social media behemoth has devised "Vistara," a chip using the open compute express link (CXL) interconnect to expand memory in server systems.

The system works by taking fourth-generation double data rate (DDR4) memory modules from decommissioned older servers and installing them into newer systems that normally use DDR5, with the Vistara ASIC exposing the older DDR4 as additional memory using CXL.

According to Meta, the Vistara "bridge" achieves "substantial gains" for diverse workloads, with up to a 25% reduction in server count for disaggregated machine learning (ML) inference; 29% reduction in average latency for distributed caches, and a 33% drop in the frequency of job failures.

Vistara validates CXL

In a research paper, authors from Meta revealed approximately 40% of the company's servers are memory-capacity bound due to the current bottleneck in memory thanks to the boom in AI demand.

With memory boasting a lifetime of of up to 10 years compared to up to five years expected from servers, the researchers sought to use CXL – defined by a "striking lack of large-scale, real-world workload experience" – without being blighted by low bandwidth, high latency, and extra computing demands associated with memory sharing across additional layers.

Vistara solves this on the hardware side, described as being optimized for DRAM reuse, power efficiency, and low latency. On the software side, Meta built an optimized solution based on transparent page placement (TPP), which separates memory pages in either fast local DRAM and slower CXL-attached memory. This helps to define the appropriate local-to-expanded memory ratio for each workload, plus automate per workload configuration, while disabling expanded memory for workloads that can't tolerate the increased latency.

Specifically, the operating system views the application-specific integrated circuit (ASIC)-connected DDR4 as a completely distinct, central processing unit (CPU)-less non-uniform memory access (NUMA) node, isolating it from the standard DRAM attached directly to the processor. Meta's platforms are designed to optimize these resources by exhausting all available local memory first, tapping into the CXL-enabled memory pool only when the workload requires the extra capacity.

Under the hood

According to Meta, its ASIC is designed to bridge DDR4 memory to host processors via a CXL 2.0/1.1-compliant peripheral component interconnect express (PCIe) Gen5 x16 interface. Driven by various custom RISC-V processors, each Vistara ASIC folds in two independent 72-bit DDR4 memory channels, supporting speeds up to 3,200 megatransfers per second (MT/s) and up to 256-gigabytes (GB) per chip with 64 GB dual-in line memory modules (DIMMs).

On the architecture side, Meta integrates its Vistara hardware into specialized machines known as MemServers, which are each anchored by a high-performance AMD Turin processor featuring 158 cores and 316 threads. To handle massive data demands, every MemServer pairs 768 GB of primary DDR5 memory with 256 GB of DDR4 via Vistara.

Vistara CXL cards meanwhile are housed in dedicated, rear-accessible compartments within the chassis.

As recently explored by this title, recycling memory to overcome the so-called RAMpocalypse has become fashionable as of late. Vast Data were arguably first out of the door with their Amplify program, which aims to replace existing storage or database software with Vast's own platform. Assessing a client's solid-state drive (SSD) estate, Vast reclaims non-volatile memory express (NVMe) drives from legacy arrays, data lakes, and existing servers.

Rivals like Vdura have argued against reclaiming flash in any form, with Vdura’s SVP of Business Operations Erik Salo warning SDxCentral that as flash reliability deteriorates as it wears.

"If you have ... half a dozen flash drives fail, your data is gone. ... I'll tell you, I would not use old flash," Salo said.

Weka Chief AI officer Val Bercovici echoed that sentiment, noting reclaimed stranded performance capacity requires serious due diligence.

“And that's not always transparent," Bercovici argued. "You also need to understand usage, wear level, and actual capability. Only if it passes all those tests is it appropriate for high-performance AI workloads versus commoditized capacity use cases.”