AI
– Getty Images

“I get asked all the time what I think about training versus inference – I'm telling you all to stop talking about training versus inference.” So declared OpenAI VP Peter Hoeschele at Oracle’s AI World 2025 event, suggesting that as AI models are now designed to run continuously, models are no longer running separately across training and inference modes.

Similarly, John Bradshaw, field CTO Cloud for EMEA at Akamai Technologies, sees a move away from AI training into AI inference, saying the rubber doesn’t necessarily have to meet the road.

“Training is great if you're in the business of training things, but it doesn't actually give you any value,” Bradshaw told SDxCentral. “It's really useful, but until you exploit what you've trained, it hasn't done anything for you.”

The CTO sees a current shift from “big, centralized AI shops” into more discrete solutions, including the domain-specific sort and small language models (SLMs), as opposed to large ones.

The signs, therefore, point to 2026 being the breakout year of AI inferencing – the process where a trained machine learning model generates predictions and outputs from new input data. This is done across two phases, with compute processing a prompt, and the model then generating output tokens.

Dell’Oro Group recently singled out inference requirements for foundational models as helping drive deployments of custom accelerators from hyperscale customers such as Google and Amazon. This saw global data center server and storage component revenue grow 40% year-over-year in the third quarter of 2025, according to Dell’Oro Group research.

Nvidia has already kicked off the new year by all-but-acquiring AI chip firm Groq to galvanize inference via hardware, alongside rumors it may acquire large language model (LLM) makers AI21 Labs for the software side of things.

Meanwhile, Ishit Vachhrajani, global head of technology, AI, and analytics, at Amazon Web Services (AWS), stressed to SDxCentral for this feature that inference engine Amazon Bedrock was already a “multibillion dollar business” for the hyperscaler.

John Bradshaw AKAMAI bnw
John Bradshaw, Akamai – John Bradshaw | Akamai

For Akamai customers, inferencing is driving hyper-personalization in e-commerce and travel, providing tailored product recommendations based on individual preferences, as well as agentic AI customer interaction.

Bradshaw also noted one company, Monks, which is automating media workflows such as real-time camera switching in sports events, with AI analyzing video feeds and selecting the most relevant angles without human intervention.

The Akamai advantage is edge based, with Bradshaw claiming it has “probably the world's largest network” at its fingertips, a network of points of presence (PoPs) bringing inference workloads geographically closer to end users than typical hyperscaler regions.

“It can shave tens-of-milliseconds off between our locations and where their users are,” Bradshaw said, noting this helps a client like Monks process images in real time for live events. “We're able to pull so much latency out of that overall transaction. Even though these GPUs are not on that race track … it's possible to pull the images back and back out again in time for it to be seamless.”

AI: bringing mojo back to the edge

IDC recently predicted AI use cases will spur edge computing spend to nearly $378 billion by 2028. As such, everyone in the cloud and network space is setting out their stall and keen to talk up the inference use case, implying a diverse market where hyperscalers are not necessarily the first – or even most apt – port of call.

Ciena, for example, has stepped up to the plate with recent releases such as the Waveserver E-Series, offering transport capabilities in a small form factor. Vimal Pindoria, Ciena’s European director for routing and switching business development, said the Ciena vision is preparing business “for what could come by enabling an adaptive network.” This means scaling last‑mile links from 10 Gb/s up to the 100 Gb/s and higher expected with AI workloads, adding network capacity from the enterprise to its first point of aggregation and to the cloud.

The hardware is only part of the equation, though. Renen Hallak, CEO of AI storage firm Vast Data, stressed that inference creates tough data requirements at the edge, as it’s a workload that “needs to be up 100% of the time, because it’s a production workload versus training which can go down and nobody will notice.”

As a solution, Hallak argues against a shared‑nothing design where each server or processor in a cluster is fully independent, instead advocating “a shared-everything way” so that as you add nodes you get more performance, more resilience, and more capacity efficiency.

Nutanix is thinking along similar lines, with its Chief AI Officer Debo Dutta positioning its Enterprise AI product as a shared inference infrastructure that timeshares scarce accelerators across multiple apps.

“When you have distributed infrastructure, what you need is the operating system or the software to manage the complexity and lifecycle," Dutta explained. "You have to abstract out all the hardware complexity and all the lower level complexity to give very simple cloud-like constructs, but under enterprise IT’s control.”

Another vendor setting out its edge stall for the inference revolution is Cisco. The network giant has recently been touting its Unified Edge system, which combines compute and GPU resources with Cisco’s networking and SD-WAN technologies in one platform designed to run inferencing closer to where data is generated.

Jeremy Foster, SVP and GM of Cisco Compute, explained AI, specifically inference, “breaks the old cloud model” by forcing distributed architecture across edge, core, and cloud to work together.

“That is going to put pressure on a lot of things,” Foster said. “Whether that's how you’re going to secure all these different distributed environments, to how you are going to autonomously operate these environments.”

According to Foster, as AI scales into everyday use cases, infrastructure is going to require autonomous operations.

“It's not going to be like a nice to have. It's going to be really the only way to run your organization with these globally distributed systems. That's where the edge components become important to try and help enable that piece of it," Foster said. “But it's not to say that we think inferencing isn't incredibly important in the data center – but it's going to take all those things to be brought together into a big platform that you can use to be able to derive value from.”

Jeremy Foster Cisco
Jeremy Foster, Cisco – Cisco

Foster also believes AI training is not a thing of the past.

“I think training is something that's going to always have to happen," Foster said. "But from where do businesses derive value, and where does an enterprise customer want to start? You're not going to win by building an application enterprise just because you have a slightly better model, because everybody's going to have access to these models, and these models, to some extent, become commoditized.”

With Unified Edge, Cisco is targeting retail, hospitals, and manufacturing sectors that increasingly rely on agentic AI apps which are capable of making autonomous decisions at the network edge.

Foster claims these enterprises win out by tweaking their models with the right data for a competitive advantage, and by leveraging AI models to build the right products for their customers. As such, demands for AI training will increase, which in turn drives more inferencing.

One example of a Cisco client using inferencing is a manufacturer operating computer vision across various plants and thus amassing a huge amount of data which is then processed by the edge system. This data is then shared with the cloud to retrain an AI model.

“Now those machines are operating more efficiently than before," Foster said. "Things like that are going to happen more and more as we move forward, and those are the types of business opportunities that are out there for inferencing. It enables a continuously improving business process.”

This can benefit “any company or size or scale” according to Foster, and not just the Fortune 500 brigade. Interestingly, the Cisco VP noted that even bigger enterprise accounts don’t necessarily have data centers ready to deploy inference workloads.

This is where the neoclouds or AI clouds come in, a segment which has seen an 82% compound annual growth rate in revenue since 2021, according to recent JLL research. For Foster, the likes of CoreWeave, Nebius, and Crusoe can be seen as specialist firms compared to the more general service-provision of the hyperscalers.

“Some are specialized in inferencing use cases, other ones are focused on training, and you might have other ones that are focused on providing bare metal for these types of GPU systems, so a customer can do whatever they want. And then you get into cloud, where a lot of customers have a lot of their data," Foster explained. “Ultimately it's about bringing those three things together, which is the cloud sides, what an enterprise customer wants to deploy, whether that's [on-premises] or in a neocloud, or both, as well as what are they doing out in the field, if you will, with edge.”

‘Open’ AI for inferencing

The rise of neoclouds is one reminder that the AI training and inference case is not necessarily a market owned by the likes of Google Cloud and AWS.

The expensive nature of AI may be a deciding factor. Akamai’s Bradshaw claimed its cloud services are up to two-thirds cheaper than the hyperscalers when it comes to compute GPU resources. As such, there is an appeal to open cloud, and, in turn, to actual open AI (not to be confused with OpenAI).

On this front we have Akamai promising a simplified developer experience with an open source and portable cloud application platform, supporting the build out of retrieval-augmented generation (RAG) pipelines and other components into AI workflows.

As to be expected, giants like Red Hat are also touting their form of open AI with virtual LLMs (vLLMs), an open source inference platform used within Red Hat products and compatible with hyperscaler solutions.

Red Hat’s Senior Principal Technologist for AI Robbie Jerrom highlighted that with AI developing so rapidly, there’s “always a new model to try out.”

“You can run any model on any cloud or in your data center across that consistent platform, and that choice is what customers want," Jerrom said. "It's such a fast moving area. There's always something new coming in.”

Jerrom highlighted the LLM-D project, which works in conjunction with vLLM to reduce the cost of inference and improve inference performance. This is achieved with key-value cache (KV cache), a token cache that registers prior interactions with a model to avoid activating the GPU when those interactions are repeated.

“If you start doing that over time with lots and lots of user responses, you can get cache hits of up to 80-88% and that drives down cost, but more importantly, it increases the performance, and we can share that cache across multiple models," Jerrom said.

Cisco’s Foster also believes the performance aspect is more important than cost for customers, highlighting that scale is everything when it comes to pricing.

“In general, building a giant training cluster is more expensive than inferencing. I can do inferencing on small devices,” Foster said. “Working with enterprise accounts, you’re asking, ‘what do you actually need to deliver against your use case?’ What's easy to do in AI is overspend and get more capacity than what you might need right now, particularly as you're starting to figure out how these use cases are going to be deployed and then eventually scaled.”

The Amazon response

But some classic cloud conundrums may still define the AI market landscape. Nutanix’s Dutta gave a reminder that regulated customers can’t rely on public clouds, with the same guardrails needed for their AI workloads.

The AI chief posited Nutanix Enterprise AI as an inference platform similar to Amazon Bedrock, but “on your own private infrastructure for the purpose of inferencing.”

“So within a few clicks, you can deploy a model from Hugging Face or Nvidia, or your own private model, which is usually the case for many of our regulated customers, and then give you a shared inference endpoint," Dutta said.

Akamai’s Bradshaw, meanwhile, claimed customers are “very concerned” about high egress fees with the hyperscalers, and “they're not banking on those same hyperscalers to deliver their AI workloads,” instead opting for vendor neutrality.

Arguably, the egress question is an issue if Big Cloud users are storing data and running applications elsewhere and using the hyperscaler mainly as an inference endpoint or accelerator by constantly shipping data and results back out.

The Amazon response is simple: run your inference with us. AWS' Vachhrajani pointed SDxCentral to various benefits, including its custom silicon, innovations in liquid cooling, Project Rainier compute cluster, and general supply chain prowess.

Specifically, Vachhrajani touted Amazon Bedrock as on track “to be the world's biggest inference engine,” with more than 50% of all of the tokens currently generated on Bedrock running on Amazon's custom chips.

He also expects up to 90% of all of AWS workloads and spend to be inference-related.

Ishit Vachhrajani, AWS
Ishit Vachhrajani, AWS – AWS

“We expect Bedrock to be as big a business as EC2 is in the future, as well as the application layer, which is where you get to use inference and AI in how you serve the customers," Vachhrajani said.

AWS takes the view that inference will be a key building block for modern applications, similar to how having compute storage and database powers websites, ERPs, CRMs, and transactional applications.

“We are actually putting inference in that same plane,” Vachhrajani said, echoing AWS CEO Matt Garman’s recent prediction that the cost of inference will go down by 10x.

“We are at that space where the cost of intelligence is dropping and the level of intelligence is rising," Vachhrajani added. "And I think this is the sweet spot for many, many use cases to actually start leveraging AI in a cost efficient fashion.”

One customer doing so is life science firm Metagenomi, which Vachhrajani claimed brought their inference costs down by 56% when creating protein language models. Such modelling requires huge and thus expensive amounts of inferencing, but the Amazon hardware and software stack saw savings, with Vachhrajani also pointing to the networking backbone.

“Because the latency, especially when you think about inference, becomes very critical, and this is why we have innovated in our networking,” Vachhrajani said.

The neocloud proposition

The AWS head is also not fazed by neoclouds, echoing Cisco’s views that they currently only serve niche workloads.

“We don't believe [inference] is going to be an isolated application. At the end of the day, intelligence and inference and AI will be part of your applications through and through … CIOs and CTOs don't want to manage additional complexity of moving data, building integration, building the pipe. They want to focus on delivering the better outcome for the customers," Vachhrajani said. “I think this is where having not just a particular solution in one space benefits, because enterprise workloads are diverse and complex, but something that meets the needs of customers through and through is a critical differentiator that AWS brings to the table.”

Some industry insiders have also claimed AI-native neocloud firms spend more money on GPUs than they do on latency and networking, affecting their rentability.

One AI cloud disputing this assertion is CoreWeave. Corey Sanders, SVP of product at the firm, underlined how CoreWeave’s value to the inference question is in its combination of software and infrastructure around its GPU provision.

“When you're running an inferencing workload that spans 1,000 nodes, and one is going slower than the others, you want to know why,” Sanders said.

As such, the neocloud offers observability across large distributed clusters, with a new telemetry relay feature collecting observability data to share with the client's monitoring system of choice, keeping an eye on factors such as a GPU’s reliability and health.

Corey Sanders, Coreweave
Corey Sanders, CoreWeave – CoreWeave

Where competitors have focused on capacity issues when it comes to AI, Sanders goes a step further by noting customers are more likely to face issues on the model level.

“On the customer application side, it's sort of hard to build the models, it's hard to do your testing, how to do your evaluation, how to put the right guardrails in place to make sure that when you deploy it's doing what you want it to do," Sanders said. "The inferencing world has been a bit slow to sort of build this into their applications.”

As such, CoreWeave offers a serverless reinforcement learning (RL) focused on making all those AI agents out there more reliable while learning correctly as they progress business actions.

“Instead of every issue having you go back and tweak the prompt, with reinforcement learning you can sort of make this learning as you go and make your model and inferencing better and better and better," Sanders said. "I expect this to become very commonplace over the next year.”

The CoreWeave customer set is described as a healthy mix of the enterprises, which hyperscalers tend to court, as well as labs and digital-native customers who have built their solutions in the cloud and are now pivoting toward AI.

In addition, Sanders noted that opposed to niche workloads, 50% of current work being done on CoreWeave is AI inferencing, with the technology becoming more mission-critical to its customer set.

Storage is key

Another way CoreWeave is extending beyond the hardware with software to improve rentability is its multiregion, multicloud Local Object Transport Accelerator (LOTA) offering to power AI object storage as inference fuel.

“It is GPU aware and understanding of the customer environment so that we can cache what they'll need for their workload and even share it across GPUs to make sure that we're delivering the most optimized amount of data going into the GPUs,” according to Sanders.

Underneath something like LOTA is an underlying storage solution, provided by the likes of CoreWeave partners such as Vast Data.

Presenting itself as an AI storage vendor of choice for both hyperscale and neocloud platforms, Vast has recently announced partnerships with Google Cloud and Microsoft Azure.

For Vast CEO Hallak, hyperscalers are “now trying to make up for lost time,” encumbered by legacy whereas AI clouds “have found a way to build this stack faster from the ground up, because they didn't have that legacy, and when they look for the software stack they come to us because we built the software stack for these AI workloads.”

“For inference … [and] wherever you need large GPU clusters, the old hyperscalers are not as capable as these neoclouds," Hallak added in a media briefing attended by SDxCentral last year. “So now we're starting to see the big hyperscalers wanting to adopt our stack as well."

Renen Hallak, VAST
Renen Hallak, Vast Data – VAST Data

The benefits are mutual, with Hallak saying a vast amount of Vast customers are using inference and other AI workloads through hyperscalers in their cloud stacks, highlighting that AI work is not yet the exclusive domain of the neoclouds.

“They want to get the same experience that they're getting [on-premises] and in the AI clouds as part of their hyperscale deployments," Hallak said. "And they're asking the hyperscalers for the same and when the customers demand it, then we have to do it, and they have to do it, and so we start building."

Hallak argued that a full AI operating system designed for AI workloads rather than generic enterprise applications will better suit enterprises in five to 10 years when fine‑tuning via reinforcement loops and inference will be “99% of what gets done.”

This will occur as AI usage moves from simple question‑and‑answer interactions to multihop reasoning where agents must “cross correlate and validate things,” making inference “even more computationally intensive.”

“As we shift also from training to inference, the weight shifts from compute centric to data centric, because inference requires us to look at unstructured data and cross correlated with database abilities," Hallak said.

An optimized data platform for both unstructured and structured data may be the solution, with Hallak arguing that you can’t just add GPUs, but instead must redesign the data layer to keep them fed. Such a unified data layer keeps the GPUs compute‑bound rather than input/ouput‑bound as inference scales up.

Hallak is clearly already thinking beyond the current era of chatbots and everyday software agents to what the future holds for inference, referring back to the use of computer vision as cited by Cisco.

“As you have these agents, they're sensing things, they're seeing things through video cameras, they're hearing things through microphones, and then they need to infer and understand what they're seeing and hearing," Hallak said. "And so that's thought, a structured piece of information, and then they need to cross-correlate that with something they saw a month ago or a year ago, and that's their memories or old thoughts that they had before.”