A majority of organizations are planning to expand their artificial intelligence (AI) infrastructure in 2024, but moving too quickly without appropriately prioritizing generative AI (genAI) use cases is a major concern for IT leaders, according to a new report from the AI Infrastructure Alliance (AIIA).
The survey reached AI/ML (machine learning) and other technology leaders at 1,000 global organizations and found that an overwhelming 96% of respondents are working to scale their AI compute infrastructure, capacity and investments in 2024. More than 50% of respondents plan to use large language models like LLaMA, and 26% plan to use embedded models in their commercial AI deployments.
The availability, cost and architecture design of this high-performance infrastructure, however, all present substantial challenges.
IT leaders are well aware of compute limitations like GPU availability and overall cost, but most pressing is the concern that these organizations are deploying genAI applications too quickly and missing important considerations – like prioritizing the wrong business use cases.
The second largest concern in deploying genAI applications, ironically, is moving too slowly, whether that’s due to lack of execution ability or ambiguity among leadership. IT leaders are caught between the pressure to innovate and the dangers of making mistakes.
To that point, companies are looking for clarity, peer reviews and industry benchmarks on the various AI platforms available today, according to the report.
When evaluating various AI infrastructure against business use cases, IT and AI leaders also consider factors that affect the total cost of ownership (TCO), including compute, scheduling, latency and power efficiency. “Only then can [IT leaders] be confident in accurately predicting and forecasting the TCO for genAI in their organization,” according to the AIIA.
Addressing GPU scarcity The report demonstrates a meaningful shift in AI hardware usage and a growing demand for financially efficient compute for AI inferencing. To confront GPU scarcity in 2024, 52% of respondents are actively searching for cost-effective alternatives to GPUs for AI inferencing. Just 27% are looking for wallet-friendly alternatives to GPUs for AI training, according to the report.
These findings signal “a need for highly performant, cost-effective ways to optimize GPU utilization, or find alternatives to GPUs,” ClearML CMO Noam Harel said.
“Businesses are actively looking for new, cost-effective options for inference compute,” FuriosaAI CEO June Paik added. Both ClearML and FuriosaAI are part of the AIIA.
The survey also found idle GPU resources are a major culprit of runaway compute costs. While 78% of respondents are using more than 50% of their total GPU resources during peak periods, just 7% report their GPU infrastructure reaches more than 85% utilization during peak periods. This further indicates a need to improve management of existing compute resources and expand infrastructure with alternatives.
Comments