With all the talk about AI transforming the data center, a perhaps overlooked element is the essential nature of optical fiber in next-generation networks. A failure in just one network link can disrupt or even halt operations within an AI cluster, reducing performance and therefore jeopardizing substantial AI investments.
Bill Gartner, SVP and GM for optical systems at Cisco’s Optics Group, reminded SDxCentral that such failures were once a minor concern in traditional IP networks thanks to transmission control protocol/internet protocol (TCP/IP) tending to bridge the gap by retransmitting a packet.
“But the insight that we now have, having talked with many of our hyperscaler customers, is that AI infrastructure is much, much more sensitive to that problem to an error burst," Gartner said.
This is down to how GPUs run in parallel, with one error in the chain forcing each processor to stop operations, back up to a checkpoint, and restart.
“We've got some data from third parties that suggests that, with [OpenAI’s] GPT-3, there were basically like 10,000 GPU infrastructures, and GPT-4 is now going into tens-of-thousands, and with GPT-5, we'll see 100,000 GPU clusters," Gartner noted. “Start doing the math, and if you have an optic with a typical five-year mean time to failure, you think that sounds okay. But in a 100,000 GPU cluster, your first failure is going to occur in 26 minutes, and so that has a big implication on the utilization of that AI infrastructure. And what that can mean is that a single link failure can reduce the overall performance of the cluster by like 40%.”
Compounding the problem is that it’s difficult for users to isolate the exact cause behind an error and determine whether the optic or fiber contamination was at fault.
“They end up just randomly replacing that optic,” Gartner asserted.
Another issue is that not all optics are built equally. According to Gartner, Cisco recently tested 20 optics between 100 Gb/s and 400 Gb/s capacity from various suppliers. Despite claims of full compliance on all optical and electrical standards, not a single one passed its qualification process.
“Customers may choose to acquire third-party optics and put them into routers and switches that they acquire from Cisco," Gartner said. "What we know is that the optics that they're acquiring are not at the same level of reliability as ours.”
The laser crunch
The bigger question is how much supply meets demand in today’s market. A recent McKinsey study warned that next-generation networking won’t be hindered by GPU development but by an optical supply shortfall, with a laser crunch throttling AI-era networking through to 2030.
According to McKinsey, production of 800 Gb/s transceivers is projected to run 40% to 60% below demand until 2027, and shortfalls of 30% to 40% are likely for 1.6 Tb/s devices through 2029.
“Laser supply, in general, has been a challenge in the industry,” Gartner concurred. “So I'd say there we have seen challenges. We do everything we can to try to get ahead of that. … Beyond that, I think we are in pretty good shape on the supply chain.”
The laser squeeze comes not from service providers but from hyperscaler demand. According to McKinsey researchers, hyperscalers will shift around 87% of their back-end optics to 800 Gb/s and greater by 2029, with 1.6 Tb/s transceivers accounting for more than 40% of this demand.
Gartner says Cisco is buoyant on optics, seeing a “dramatic” demand driven by AI and AI-driven scale across networking, with the firm meeting its billion-dollar target from last year on deals with hyperscalers across AI applications and infrastructure.
“In the first three quarters, through roughly half of that business that we had done, we achieved almost three quarters of that billion dollar target. Half of that was optics and optical," Gartner said. "So by the end of the year, we had done well over that target, and around a third to half of that was in optics and optical.”
Gartner also noted hyperscalers have a desire for smaller form factors and have thus greatly influenced the optical systems market through their adoption of pluggable technology at a “much, much faster” rate than most service providers.
“Virtually all of them are doing that at 400G. Most are moving to 800G, and most will move to 1.6T when that's available,” Gartner said, adding that around 350 service provider customers have adopted the same technology from Cisco.
“But I would say it's naturally more of a slower evolution for service providers in part because they have much more legacy to deal with [and] challenges with their operation systems, but the service providers are definitely moving in that direction," Gartner said, giving the example of Lumen, which recently adopted routed optical networking for their metro networks.
Pump up the WAN
Gartner also noted that the forward surge of AI is making hyperscalers examine infrastructure more closely, leading to the uptake of laser pumps, which are energy sources used to galvanize a laser's gain medium.
“Hyperscalers, and some application and service providers as well, are deploying so much fiber that they're actually looking at the architecture of these line systems," Gartner explained. "A simple example is historically one fiber traveled through a series of amplifiers in order to get from point A to point B, and those amplifiers have things like laser pumps in them.”
Gartner revealed Cisco is looking at how to share components such as laser pumps across multiple fibers in a single site.
“Can we somehow consolidate the size of those systems and leverage common components like laser pumps, which are part of the amplifier infrastructure to give it better economic performance? That's something that five, 10 years ago we wouldn't even have thought about doing because it just wouldn't make economic sense to try to think about, how do you deal with multiple fibers going through one site. It just wasn't a problem, but it is now.”
Comments