Data storage systems have been refined over the years to provide a stable platform onto which organizations can dump their prized data assets for future perusal, but AI is changing that dynamic due to the surge in new AI-generated data that needs to be stored, sorted, and primed to feed AI needs, changes that are upending the structure and scale of data platforms and how those platforms are being used.
While much of the AI discussion of late has been around compute resources needed to run AI data centers, many have noted a growing correlation to storage needs for dealing with the crazy amount of data being driven into and emerging from these locations.
Sam Grocott, SVP of product marketing at Dell Technologies, recently told a press briefing that this crunch was one of a handful of challenges the vendor was hearing from customers. One of the most consistent is that “data growth isn't slowing down. Storage has got to keep up, and it's got to become more and more efficient as well.”
Grocott outlined that acceleration by noting data center storage was projected to grow at a 19.5% compound annual growth rate (CAGR) to more than three zettabytes by 2029, “so data is absolutely the fuel for AI, and the density to store it and then to orchestrate it across the AI solutions is becoming more and more strategic than ever.”
IBM added more numbers to those expectations when it recently cited a Precedence Research study that predicts the global AI-powered storage market will grow from nearly $36 billion in 2025 to more than $255 billion by 2034, surging a significant 24% CAGR over that decade.
More stressing, IBM conducted a survey that found while 62% of executives expect to use AI across their organizations over the next three years, only 8% said that their IT infrastructure was up to dealing with that usage. Despite that gap, only 42% of those surveyed think their infrastructure can manage the data volumes and compute demands of advanced AI models, with just slightly more (46%) expecting that infrastructure to support real-time inferencing at scale.
Don Gentile, analyst-in-residence for storage and data resiliency at HyperFrame Research, told SDxCentral that recent research the firm has done on the topic showed a vast majority of enterprises are bogged down by growing operational complexity that is stunting their ability to deal with this challenge.
“Reality is that a lot of organizations are still held back by the complexity; they know they've got silos, they've got teams that are not talking to each other, they've got systems that they don't necessarily trust to be part of the AI pipeline, and so that is the reality,” Gentile said.
Storage system structures
Storage systems typically sit somewhere in the middle of a conventional AI infrastructure stack. This is above the base systems that include the foundation infrastructure and compute resources, but below the AI models, runtime systems, applications, and services.
That storage layer includes various mediums. There are hard disk drives (HDDs) that trade on low cost at the expense of speed; and flash (NAND) and solid-state drives (SSDs) that are higher performing at a higher cost, which some have noted can be as much as 10-times more expensive on a price-per-byte basis.
Within these general constructs are different storage means like block, object, and file storage, and also hybrid models that combine HDDs with SSDs to hasten specific tasks.
Jerome Wendt, CEO and principal analyst at DCIG, explained that these systems serve various roles in a storage environment and are being altered to deal with surging AI demands. This is through the need to have that data for inference training of AI models, which are the basis for newer agentic AI systems that can support specific use cases.
“Most object storage systems, they were designed for archive, and they started putting SSDs there to make the retrievals faster,” Wendt said. “But really, with artificial intelligence, faster speeds doesn't always help. You really need to index all the data and create this metadata database – it's called vector embeddings – and then they use that pretty much like a RAG (retrieval-augmented generation) pipeline.”
Wendt explained that this model basically means the systems have “already indexed all the data, they've already got it set up, so when you do an AI query, all the AI workloads have all that data already, all the metadata are at its disposal. … So you don't have to touch every piece of data to get all the files and objects to pull back, they're just going to touch this front-end object or metadata database to get the information they need because it's already indexed, all the data I need, and it's constantly indexing all the data that's stored on it.”
This redesign can allow these object storage systems to work faster than block-and-file systems.
“It's really the performance of the metadata database where they're really hosting all the stuff, and ultimately matters for these AI engines,” Wendt said.
And these AI engines are generating a lot of data.
Gentile explained that with traditional storage, “you have an application, you have a database, you have a user, and you have a storage system, so it's pretty much like a one-shot deal. You're going to make a request to that application, that application is going to hit the database, the database is going to turn that result back to the user. There's your traffic, there's your storage, and your network implications.”
But with agentic AI, “that magnifies this by tenfold,” Gentile added.
“You've got the user prompt into the AI situation, then you've got a retrieval system and that's going to go hit the vector database. Then this storage system has a source document or multiple source documents that it's then going to have to retrieve, bring that back, and then the GPUs are going to perform that inferencing calculation on that data,” Gentile explained. “Then there's also observability tools that are tracking that, and a governance system that's logging all of that activity as well, so that is already creating a tremendous amount of traffic, continual calls back and forth between the system, so you have increase in traffic.”
For long-standing operations, these important data pools can reside in long-forgotten systems that have yet to be tied into this AI workflow, or worse, subsumed by new data storage systems that are bought in hopes of gaining new storage capacity.
Wendt called this the “just buy more philosophy,” which organizations are going to need to ween themselves from if they want to keep a handle on costs and not overlook their standing data pools. This means having a storage administrator empowered to oversee these operations.
“That [just buy more philosophy] is starting to break and people are going to have to be much more purposeful and thoughtful about what they're doing on that front,” Wendt said. “There are companies that can help on that, but you're going to need people to manage it, someone who's very thoughtful and appropriate for the role. Otherwise, it comes down to someone in the system admin or storage admin chair saying, ‘OK, I'm now setting the storage strategy for the company because I'm the one administering this and no one's giving me any guidance on how to best manage the state.’ So you end up, by default, making policies on governance and retention and everything else.”
Gentile echoed that sentiment, noting that AI models often produce unnecessary data and there is a growing need to “de-duplicate systems. … AI, it's generating so much additional storage and so you've got to find ways to compress and de-duplicate so that you have less storage usage overall.”
Wendt added that management of these systems and decisions are paramount, especially as data amounts surge.
“If you have over a petabyte of data, you cannot run all the backups, you can't do a restore, you can't really even do a [disaster recovery] very easily with a petabyte of data,” Wendt warned. “If you're at that level that you're actually managing, you better get out of the 'just buy more' philosophy and just keep throwing more stuff at it because your infrastructure is already broken. You just don't know it and you're going to find out at the worst possible time.”
Storage cost, market uncertainties
That “just buy more” model is also running into broader cost concerns, which include the oft-discussed memory shortage that is expected to last into the foreseeable future; growing concerns over “tokenomics” and the cost of transferring AI data; and recent competitive moves.
Gentile noted that the “elephant in the room” of “supply constraints and costs” are “putting tremendous pressure on the industry, particularly the enterprises.”
“I would say the hyperscalers have locked in their contracts … so they’re going to get their GPUs,” Gentile said, adding that “the SSD vendors in particular are in a massive … 18, 24-month supply constraint issue, which you're hearing from everybody. It's just the way it is.”
Then there are concerns about the cost of AI-related data transfers. This is typically tied to token costs, which, while they have plunged over the past year, remain a considerable issue due to the growing use of that payment structure.
“Tokens are not unlimited; most people have finite amounts,” Gentile said. “The cost of tokens will go down over time. … There's some sort of asymptotic graph where there's a happy medium, there's an equilibrium where an enterprise is actually using what they can afford, and I think everybody's kind of trying to find out where that sweet spot is.”
This challenge was central to comments from Cisco Chief Product Officer Jeetu Patel, who, during that vendor’s recent Cisco Live event, expressed a need for the ecosystem to attach the right value to tokens.
“When you start thinking about where the risk lies, it would be when the cost of tokens and the value derived from the tokens actually have a distance, and that's the thing that we, as an industry, have to be really careful of is you want to make sure that you're economically generating tokens that create the necessary and desired output from an end-result perspective, on the core metrics that you, as a business, are trying to go out and measure, because that is going to be extremely important, that is in equilibrium with the cost of the tokens,” Patel explained. “If your costs go out of whack, but the benefit is not there, that's when you will actually see some pullback.”
Patel added that this is why it’s important for an organization to get an early handle on their AI usage.
“I feel like across the board right now, the first phase that you get in AI is you have to get good with using it, which means you have to get familiar with it first. That consumes tokens,” Patel said of that initial stage. “Once you get familiar, then you get good and that's when you start creating good outcomes. And once you get good, that's when you start seeing very strong accretions of value with your company.”
The market overall is “still in phase one,” Patel noted. “Uniformly, all companies haven't gotten good at this yet. We have to get good. Once you start getting good, you start seeing value and outputs get accreted in a very quantitative way.”
Some enterprises are also faced with vendor decisions that could force new data storage directions. Wendt pointed specifically to Broadcom’s ongoing VMware pricing and licensing changes that have impacted how enterprises view the vendor’s vSAN storage platform.
“There's this internal pressure just to get off of VMware, switch over to a competing solution, … and well, they often need new hardware for that,” Wendt said of that challenge. “Can you
still use your VMware hardware that you're hosting, can you reuse that for this new solution? Maybe yes, maybe no, but I think that’s also driving some new demand. Sometimes they have to run it both in parallel for a while as they have to keep the old stuff running, and then you get everything transitioned over. It's not like you just unplug one and then plug in the new stuff and it's running the next day. So just a lot of crazy forces going on here in the market right now.”
That craziness is not expected to subside anytime soon, which will continue to put more pressure on both sides of the data storage pendulum to figure out ways to manage this AI-fueled challenge.
“This is very scary territory for people to suddenly say, 'oh, yeah, throw it into our AI pipeline, see how it goes,’” Gentile added on this challenge. “Many are not going to take that risk.
Comments