It’s a cliche to say that “data is the new oil,” but what does that mean for enterprise AI? How has it evolved in the past 12 months?

Let’s start by debunking the “data is the new oil” analogy. While oil is a fungible and finite commodity, data is mostly unique and can be infinitely created.

What is true is that data, particularly the enterprise’s proprietary data, is the fundamental source for the customization of AI models for specific companies, industries, and use cases.

Most enterprises are in the process of planning AI-enabled applications, and many have successfully deployed some into production. The process of moving proof-of-concept AI projects into full deployment continues to be an area where many projects face significant barriers and fail to make the transition.

These include elements such as significant project costs – including the AI development and deployment infrastructure, project goal alignment with stakeholders, and scarcity of AI development talent.

The enterprise data lake

The enterprise data lake – where relevant organizational data is collected from siloed applications, shared drives, and log data – is the most common method to utilize such data to build AI learning models.

Identifying, aggregating, extracting, normalizing, and other data ingestion tasks, while often the most time consuming and labor-intensive part of the AI development process, are essential to create an accurate AI model. Modern data lakes (as compared with legacy Hadoop-based data lakes) are based on object storage software, which is often stored on disk-based storage servers for cost efficiency.

Once ingested, the data is then used to fine-tune commercial or open-source large language models (LLM), as in the case of generative AI applications. Rather than create a completely new LLM from scratch, a pre-built language model which incorporates commonly available domain knowledge but requires additional training using enterprise-specific data.

This enterprise model fine-tuning stage requires dedicated GPU and storage infrastructure, and is a continuous process as new data is created. The result is a customized LLM, which contains the company’s specific information to generate responses.

It’s often not feasible to retrain enterprise LLMs every time new data is created, especially in the case of real-time data such as financial market information, news, and other temporal data. In these cases, retrieval-augmented generation (RAG) has become a popular technique. It appends contextually relevant information to the input query which is used to augment the original query.

The retrieval phase searches a vector database for similar information where the relevant information has been previously stored in the form of vector embeddings – numerical representations of the data – and then combines it with the tokenized original query as input into the enterprise LLM.

This method produces more relevant responses and reduces hallucinations. The vector database used in RAG is a data store which can be implemented as file or object storage.

Large-scale inference

One change from 2024 is the implementation of large-scale, high-volume inference as a part of agentic AI workflows, which combines a series of reasoning or goal-seeking series of AI agents. This high-volume AI inference workload processes thousands of queries per second, requiring more optimized efficiency in data processing.

One optimization that is starting to become adopted is known as a disaggregated inference process, which separates the two phases of processing an inference query. First, the prefill phase tokenizes input query, and the second decode phase outputs the AI model response. By dedicating separate GPUs resources for each phase, the overall inference throughput can be improved.

In the decode phase, a further optimization is to store the results of previously processed queries to look up the results when the same token pattern is presented. The key-value (KV) cache stores these previous results in multiple tiers – from the very fast but small-scale GPU memory, to larger, local in-system NVMe storage, and then further to shared large-scale network storage.

This KV cache can grow to multiple petabytes, using shared NVMe-based file or object storage for these intermediate token results. Processing bottlenecks from continuous recomputation of the same query can be eliminated. By referring to results stored in the KV cache, the overall inference performance is greatly increased.

Changes in the storage infrastructure

Storage and data management continue to be an integral part of enterprise AI infrastructure. This includes storage servers, networking, disk, and flash media, which all build a foundation for retaining enterprise data in a persistent and protected environment. Both disk and flash-based storage are used in enterprise AI infrastructure with tradeoffs in cost and performance.

Data management refers to the storage management software used to maintain and update digital information. This can be block-based, file-based, or object-based, and while each storage access method has a role in the enterprise AI infrastructure, they again offer tradeoffs in performance and flexibility to accommodate fixed or variable sized data.

A new element of data management is the introduction of data orchestration, which adds intelligent and automated workflows to the data management platforms.

Data having gravity is a common metaphor referring to the difficulty in moving large data sets to different computing resources. As enterprise data stores grow, the concept of data gravity will drive more of the AI computing workload to be done “in place.” This means that the computational resources will come to the data or be incorporated into the data management platforms rather than moving the data to the compute resources.

Storage ecosystems need to be a solution for AI

The role of data in enterprise AI includes the data lake for aggregating the enterprise data and using it for enterprise-specific AI model training fine-tuning. As AI becomes more prevalent in corporate environments, businesses need to adopt tools like RAG inference, which contains a vector database to enable the fast lookup of enterprise specific information related to their AI queries.

Another new trend is the need for large-scale inference and the implementation of disaggregated inference processing, which contains the KV cache data, primarily stored in flash-based network storage.

With the state-of-the-art enterprise AI infrastructure and processes continuing to evolve and improve, the companies implementing these projects need to develop infrastructure that is flexible, reconfigurable, and able to support new AI deployment methods developed in the future. However, the fundamental storage infrastructure and management of enterprise data will always be reusable.