Perplexity has unveiled research on leveraging older Nvidia GPUs for large-scale AI model execution.
Titled RDMA Point-to-Point Communication for LLM Systems, the paper examines how to run dense mixture of experts models (MoE) – think of GPT-5, DeepSeek, and Kimi K2 from Moonshot AI – while avoiding memory and network latency issues.
Such models can run seamlessly on AI-native GB200 or GB300 NVL72 rack systems from Nvidia, but face more problems on eight-GPU systems based on H100 or H200 GPUs. This is down to there not being enough capacity for the model’s short-term memory, known as key-value caches, thereby stymying large-scale deployments.
As a solution, Deepseek has developed an expert parallelism framework, DeepEP, using Nvidia's ConnectX network interface cards (NICs), engineered to reduce performance overhead when executing its V3 and R1 models across multiple H800 systems.
But a problem lies in a lack of primitive commands such as Send, Receive, and Write within large language model (LLM) architectures.
“High-performance computing has long used primitives (Send, Recv, Write) built on remote direct memory access (RDMA) for flexible low-latency high-bandwidth transfers … yet such primitives are rarely available in LLM frameworks. The key barrier is hardware diversity without uniform abstraction.”
As an example, Nvidia's ConnectX NICs aren’t universally used; Amazon Web Services (AWS), for example, uses its own networking protocol in Elastic Fabric Adapter (EFA).
Like the Nvidia NICs, EFA can handle 400 Gb/s of aggregate bandwidth. But it doesn’t support GPUDirect Async, which enables NICs to communicate directly with GPUs without involving the host CPU. As a result, EFA experiences higher latency in some workloads because data must first pass through the CPU.
Perplexity’s researchers found EFA requires larger messages to saturate, leading to a gap in performance between it and ConnectX, CX-7 specifically, as observed on MoE message routing in its tests.
This tallies with research from SemiAnalysis, which stated it was “not convinced” that EFA-based NICs were better than Nvidia's on either performance or user experience.
The tenacity of TransferEngine
Perplexity’s solution, TransferEngine, is made up of optimized kernels to manage GPU-to-GPU communication, and reportedly delivers lower latency than DeepSeek’s DeepEP running on ConnectX-7 NICs.
"TransferEngine is the foundational library that enables efficient RDMA-based point-to-point communication, abstracting heterogeneous hardware under a simple protocol ... [it] exposes a minimal API that abstracts away the complexity of the underlying RDMA interfaces. It supports multiple interfaces, including EFA with its scalable reliable datagram (SRD) protocol and a multitude of NICs," wrote the researchers.
The AI firm benchmarked the kernels on its in-house inference engine using both DeepSeek V3 and Kimi K2, deployed across multiple instances of Elastic Compute Cloud (EC2), powered by AWS H200 and EFA handling interconnections.
As a comparison, V3 has a total of 700 billion parameters, activating 21 billion parameters in an MoE setup for each token during inference.
Kimi K2, meanwhile, holds 1 trillion parameters with 32 billion parameters activated on any single inference. OpenAI’s GPT-4 model, as a reference point, is thought to hold a total of 1.76 trillion parameters.
The V3 tests compared a single eight-GPU setup to multi-instance configurations with 16 GPUs at two instances of AWS H200 p5en, and 32 GPUs at four instances. Performance stayed relatively stable at low and high batch sizes, but researchers claimed the multi-node setups with greater expert parallelism showed improved performance at medium batch sizes.
The same results were seen with Kimi K2, which, like Deepseek, originates from China, and represents the biggest LLM to date from Beijing-based Moonshot AI.
“Existing RDMA solutions for LLM systems suffer from vendor lock-in, with no viable implementations on custom cloud hardware such as AWS EFA. TransferEngine addresses this by identifying common functionality across heterogeneous RDMA hardware. By layering a reliable abstraction without ordering guarantees over the underlying protocols, we transparently extended support to multiple RDMA NICs, with a particular focus on EFA and ConnectX,” the researchers concluded.
Perplexity’s testing of Chinese models is understandable considering the current bans on Nvidia’s most advanced Blackwell chips to the Middle Kingdom.
Nvidia CEO Jensen Huang recently said he hoped the company would be able to sell the chips to Chinese customers “someday.”
President Donald Trump disagreed, saying this week that only U.S. customers should have access to Nvidia’s most advanced Blackwell chips.
Comments