AI networking panel
– Dan Meyer

LONDON – A stacked panel at this week’s Xcelerated Compute event was unanimous on one of the biggest issues facing the growing AI networking bottleneck challenge: reality. Though they also provided a reality-based outlet valve on how to take on that mortal hurdle: communications.

Joe Marsella, VP of portfolio management at Ciena, highlighted this challenge, noting the networking ecosystem was at a “weird moment where there’s two competing themes happening.”

“We're moving faster than we've ever moved before as an industry,” Marsella elaborated. “I mean, people are making decisions in months that used to take years. So there's an expectation that things will happen tomorrow or the next day. The flip side of that is … I've been doing this for 30 years and we've never seen demand like we're having today as an industry across the world, and that's creating, together with everything else in the world, scarcity. So on one hand we're wanting to move a thousand miles an hour. On the other hand, there's so much demand that it's difficult to meet at that pace.”

This in of itself is nothing new. Technology has often outrun the infrastructure and supply chain to fully support that potential speed, but the current speed of AI-related development is butting up especially hard against the real-world physical challenges of networking support, pushing the need for a new level of planning that connectivity path.

“There needs to be an element of planning and not an expectation that just because you want to provide this connectivity tomorrow,” Marsella said. “There's an element to fiber build, there's an element of equipment manufacturability that needs to be considered in that, and therefore planning is important.”

That planning reality is also shaped by how networks to this point have been constructed. Andy Linham, principal strategy manager at Vodafone Business, explained that this architecture has been the same for “40, 50 years … and may have added an extra layer here or there,” but more recent moves to “make a software-defined network” has added some flexibility.

“I think the benefits we have now are that because we've made that shift from big monolithic networks through to kind of more modular software and hardware overlay and underlay, we do have more flexibility built into them, so the whole network-as-a-service paradigm is not just about having a commercial consumption-based models, It's about actually, do I have access to chains that I need to make them different, to make them faster, to make them slower, to make them connect to different clouds, whatever it happens to be. We're going to have to use all of that flexibility to build out,” Linham said.

That means networks are no longer just about capacity, but, Linham said, “about where the AI services need to go.”

“There's a much wider range of places that AI could go compared to a human because it will have this access to far more data that we could ever access,” Linham added. “I think it's about architecting for reach and modularity and being able to react much, much faster to the changes as they come, because, like I said, I'm not confident in which change drives which requirement in which order of priority. So … flexibility is the key thing for me.”

And that flexibility needs to be embedded early in the process.

“I think it's about working with the actual consumers and the end users that are going to be taking those services that they develop, so if I'm building a data center with 20 megawatts capacity, having a view as to who is going to be my tenant for that capacity or having what sort of network that I need,” Linham said of this work. “The overall kind of business case planning, the more of that that we can see, the more of that we can help with, the better prepared we can be.”

Did someone say MEC?

Another benefit of this flexibility and reality is that the long-simmering mobile edge compute (MEC) model could finally become real. Linham said that “we've been looking at edge compute for a long time as an industry, and we just haven't found the use cases yet,” but “I think AI is going to change that.”

“I've talked to quite a lot of the neoclouds, and they are putting a lot of their well, not betting on it, but they're basically, they think one of their use cases is going to be wearable devices, the glasses or … all the other types of devices, they think that's going to mandate a better level of performance and a more distributed compute platform,” Linham said.

The makeup of that distributed compute platform remains in flux. Early MEC efforts pointed toward pushing these platforms as close to the end-user as possible, perhaps taking advantage of the diverse connectivity footprint of established telecom operators. But that model has been delayed due to overall costs and lack of revenue-generating use cases.

Thus, there is a more recent push toward metro-level distribution that could still take advantage of already constructed and connected locations that can serve a larger area and thus larger a radius of potential AI use cases.

“Based on the current trend where enterprises are starting to adopt rich data ingestion, it's looking like for us it's going to be around the metro and make sure that all of the data centers within a metro are well connected and well serviced with the fabric that would be high performance,” London Internet Exchange (LINX) CTO and Executive Director Richard Petrie said.

This echoed recent comments from Verizon CEO Dan Schulman, who recently teased upcoming announcements that would include support for inference computing at edge locations in support of applications like robotics, gaming, autonomous driving, and remote surgery. Verizon’s support angle comes from the ongoing decommissioning of thousands of copper-stuffed central offices that used to support legacy telecom services.

“We're already signing up deals where people are coming in … we'll talk about some of those at our next earnings call, but that ability to take advantage of what were decommissioned central offices and now can be inference, edge computing, is tremendous as well,” Schulman said of that opportunity during a recent investor conference. “It just opens up a brand new vector of growth for us on top of a growing connectivity business. We'll be more and more specific about that, but you'll start to see revenues, and noticeable revenues, beginning next year on that.”

Panel members expect that revenue opportunity to quickly trickle down to infrastructure builds, which will need to be enhanced in order to deal with this new market.

“If we start getting distributed compute, we're going to have more smaller data centers that are perhaps using, again, not GPUs but more specialized inferencing chips that are very costly and power optimized for those use cases,” Marsella said. “That means you're going to have more meshy type networks. For years, we've survived, I think, as an industry building ring-based topologies … but I think this is going to drive an even more meshy ,spider-based pathology than what we've got.”