Arrcus launched a new network fabric layer targeted at potential traffic bottlenecks caused by the growing use of AI inferencing services.
The Arrcus Inference Network Fabric (AINF) is designed to steer network traffic across distributed inference nodes, caches, and data centers. It includes query-based inference routing with policy management, interconnect routers, and edge networking, and uses a policy abstraction layer that can translate inferencing application intent to the underlying infrastructure to reduce complexity.
The platform is designed to integrate with inference frameworks like vLLM, SGLang, and Triton, and can be deployed and managed using a Kubernetes-based orchestration system.
Arrcus CEO Shakar Ayyar explained that inferencing equipment is typically deployed in distributed clusters, which can task a network’s latency and availability performance. The system allows operators to define business policies like latency targets, data sovereignty boundaries, model preferences, or power consumption goals. It can then steer AI inferencing-related traffic toward that defined criteria.
“It is now very important to realize that those inferencing nodes have a lot more diversity, a lot more distributed nature than what you are used to on the training side,” Ayyar told SDxCentral. “As a result, you can't just use a spray-and-pray approach for what you want to do at the inferencing node. You have to really understand what the application is, what the requirements of the applications are, and how that computing is actually going to be done at those inferencing nodes, and moreover, understand things like power, capacity, and the availability of power and the latency requirements.”
Ayyar noted that this can be done at an algorithmic level or down at the infrastructure level.
“We are proposing and recognizing that you can do smart things, particularly at the network steering level, start informing the policies at these routers to go in and say, here's what the application requirements are, and therefore here's how traffic needs to be steered from this point to this point,” Ayyar explained. “The richness in that policy, and being able to, therefore, then communicate that policy, and then translate that from the upper level [large language model] infrastructure or framework to the lower level, underlying infrastructure framework. That's where AINF is now essentially an innovative step forward.”
Those inferencing needs are expected to surge alongside the growing use of agentic and physical AI systems. The former has been tied to sky-high network traffic generation forecasts, while the latter has increasingly gained attention for its connection to the generation and movement of financial-related tokens, something Ayyar said AINF can help hasten.
Tokens are units of data that are typically a piece of a word, or in the case of non-text modalities, a portion of an image or an audio clip. Tokens have now become the de facto billing metric for AI services, with enterprises paying by the number of tokens that their AI consumes.
Ayyar added that AINF supports this “first-to-token” monetization route in a secure and manageable manner.
“We’re helping them manage their AI traffic by showing them how we can address the complexity of policy requirements at that network level without exposing all of that to the network operator,” Ayyar said. “We take that complexity and make sure that we can translate all those policies to effective routing policies and then handle that at different points in the network.”
Arrcus claims its ‘conquering market share from competitors’
The platform launch comes on the heels of robust growth for the vendor.
Arrcus claims it posted a record three-times increase in bookings growth last year, which came on the back of a doubling in bookings growth in 2024. However, it should be noted that the vendor did claim a three-times increase in bookings for 2021.
Ayyar said the latest boost came from “conquering market share from competitors.”
“There's no question about it. And, again, in this context, competitors are the incumbents,” Ayyar said, later noting those incumbents are rivals like Cisco, Juniper Networks, Arista, and Huawei. “It's really not about other startups or anyone else doing something similar. And to be fair, we're seeing tremendous amount of growth, and we're taking share away from competitors, but the market size in total is immense.”
Japan-based telecom giant SoftBank late last year tapped Arrcus, Broadcom, and VMware to develop a segment routing IPv6 mobile plane (SRv6 MUP) running in a live 5G-powered fixed-wireless access to provide “limited services.” That move continued long-standing work between Arrcus and SoftBank.
Comments