Meta has open-sourced CTran, the tech giant’s custom transport stack used to perform in-house optimizations.
Detailed in a PyTorch blog post, first picked up by SemiAnalysis, CTran contains multiple transport types, including Nvidia’s NVLink as well as InfiniBand, RoCE, and TCP. It provides the foundational transport layer supporting communication methods, including GPU-to-GPU transfers within nodes to cross-node RDMA operations.
CTran is designed to work in tandem with another newly open-sourced system from Meta: NCCLX, a backend system that’s an effective extension of NCCL, Nvidia’s collective communications library. Both serve as backends for torchcomms, Meta's new experimental communication API designed for PyTorch Distributed training.
Meta uses NCCLX to support large-scale AI training and inference workloads, having used it during the development of both its Llama 3 and Llama 4 foundation models.
Both can be used to scale distributed training workloads to over 100,000 GPUs, far beyond the capabilities of traditional communication libraries. The tech giant’s post suggests all of its generative AI services are backed by NCCLX, with CTran used as the underlying transport layer to make it work.
In addition to supporting multiple transport types when running NCCLX, Meta has built a fault-tolerant backend on top of CTran that provides failure detection, timeouts, error recovery, and safe reconfiguration after errors.
“Our goal is to build a flexible, extensible foundation that enables developers and researchers to move faster, scale further, and target a wider variety of hardware,” Meta’s post reads.
According to SemiAnalysis, Meta’s open-sourcing of CTran and NCCLX comes as yet another challenge to Nvidia’s collective library crown.
In addition to Meta’s newly opened sourced offerings, other competitors lining up to take on Nvidia’s NCCL include DeepEP from Chinese AI startup DeekSeek, and AMD’s Modular RDMA Interface (MoRI).
“Nvidia continues to be the leader in collective libraries, but Jensen [Huang] must not take it for granted given the heavily increased competition in the open source collective communication space,” the analyst firm wrote. “Just like how TRTLLM moved to a GitHub-first development when facing heavy competition from SGLang/vLLM, [Nvidia] should seriously consider moving NCCL to a GitHub-first open development model due to the competition in the collective front too.”
Meta’s decision to develop its own custom transport stack and communication API comes as the tech giant is making strides in developing its own in-house architecture to power its AI efforts.
Earlier this month, we got a glimpse at Non-Scheduled Fabric (NSF), Meta’s latest networking backbone for its largest AI clusters. Detailed during the Open Compute Project Foundation (OCP)’s Global Summit 2025, NSF is set to support Prometheus, Meta’s upcoming multi-gigawatt clusters scheduled to come online in 2026.
Its in-house efforts also include custom Ethernet switches, the Minipack3N. Based on Nvidia’s Spectrum-4 switching ASIC, Meta uses the switch to support both frontend and backend data center fabrics.
Meta’s custom engineering extends to chips, with its staff working to create their own training chips to handle AI tasks. Under the working name Meta Training and Inference Accelerator (MTIA), the chips are reportedly being manufactured by TSMC, with initial testing believed to have taken place back in March.
Comments