Power/Profit by Clarote & AI4Media
– Clarote & AI4Media / Better Images of AI / CC BY 4.0

Research powerhouse Gartner claimed that by 2030, large language model (LLM) training will cost 90% less than it did last year – but overall inference costs are expected to increase.

Gartner’s research focused on models one-trillion parameters in size – the earliest of which were developed in 2022 prior to ChatGPT’s breakout a year later – while taking into account AI tokens – the units of data as processed by generative AI models – sized at 3.5 bytes or approximately four characters of data.

The model’s results were divided into two categories of semiconductor scenarios: frontier scenarios, where processing is modeled using representations of leading-edge chips; and legacy blend scenarios, where processing is based on a representative mix of existing semiconductors aligned with Gartner’s forecast benchmarks.

“These cost improvements will be driven by a combination of semiconductor and infrastructure efficiency improvements, model design innovations, higher chip utilization, increased use of inference-specialized silicon, and application of edge devices for specific use cases,” Will Sommer, senior director analyst at Gartner, explained.

Sommer noted that although lower token unit costs will support more advanced generative AI capabilities, these improvements are expected to trigger a disproportionately higher demand for tokens. As token usage grows more rapidly than token prices decline, overall inference costs are projected to rise.

Savings on token costs will therefore not be fully passed on to enterprise customers, especially as businesses demand agentic-led frontier intelligence, with such models requiring up to 30-times more tokens per task than a standard generative AI chatbot.

“Chief product officers (CPOs) should not confuse the deflation of commodity tokens with the democratization of frontier reasoning,” Sommer added. “As commoditized intelligence trends toward near-zero cost, the compute and systems needed to support advanced reasoning remain scarce. CPOs who mask architectural inefficiencies with cheap tokens today will find agentic scale elusive tomorrow.”

Gartner's forecast came after LL-D, an open standard for AI inference backed by Google Cloud, IBM, Red Hat, and Nvidia, was given to the Linux Foundation this week.

The standard is based on a pre-integrated, Kubernetes-native distributed inference framework, utilizing elements such as key-value cache (KV cache), a token cache that registers prior interactions with a model to avoid activating the GPU when such interactions are repeated, saving costs in the process.

As covered in the inaugural issue of SDxCentral magazine this week, the likes of Red Hat confirmed its use of LLM-D to reduce the cost of inference and improve inference performance as using KV cache.

“If you start doing that over time with lots and lots of user responses, you can get cache hits of up to 80-88% and that drives down cost, but more importantly it increases the performance and we can share that cache across multiple models," Robbie Jerrom, senior principal technologist for AI at Red Hat, told SDxCentral.

Ishit Vachhrajani, global head of technology, AI, and analytics, at Amazon Web Services (AWS), added his belief that the cost of inference is to drop by 10-times, and while not giving a specific timeframe, the exec claimed the change has already begun in today's ongoing AI frenzy.

“We are at that space where the cost of intelligence is dropping and the level of intelligence is rising," Vachhrajani said. "And I think this is the sweet spot for many, many use cases to actually start leveraging AI in a cost efficient fashion.”

On the topic of hyperscalers, an advanced compression algorithm from Google hit recent headlines over claims it can reduce KV cache memory by “at least 6x.” Dubbed TurboQuant, the solution essentially compresses AI models while maintaining their core makeup, minus the need for preprocessing or specific calibration data.