AI inference platform FriendliAI unveiled a new offering designed to help GPU cloud operators monetize idle and underutilized capacity
Friendli InferenceSense looks to fill gaps between training and production workloads with paid AI inference jobs, with the promise of shared revenues to boot.
Claiming every idle-GPU hour results in lost margin for enterprises, FriendliAI pointed to the massive capital outlay for GPU infrastructure, with a single H100 currently renting for around $2 per hour, and an 8-GPU node at up to $20 per hour. With training and inference working on a burst-by-burst basis, hardware sits idle until the next run, meaning no fleet is at 100% utilization, it claimed, even for neoclouds at maximum capacity.
Dubbed as an AdSense of sorts for GPUs, the InferenceSense service is said to detect idle GPU capacity in a user’s infrastructure and fill it with monetizable AI inference workloads. Specifically, it spins up secured, fully-isolated containers that serve paid inference workloads, while the inference engine maximizes token throughput per GPU-hour.
When a user needs the GPUs back for their own workloads, the service is said to preempt immediately, prioritizing said jobs above others. Operators retain full control by choosing which nodes participate, as well as setting time-of-day schedules and defining exactly how much spare capacity InferenceSense may use.
DeepSeek, Qwen, Kimi, GLM, and MiniMax are among open weight AI models included in the service, with their workloads dispatched to partner hardware on an automatic basis.
FrendliAI added that token revenue generated on those GPUs is shared between it and the operator, without upfront fees and minimum commitments. Under the model, GPU capacity becomes a token-based marketplace, with neoclouds able to monetize inference demand in a more direct and profitable manner.
Elastic elan
In a recent interview with SDxCentral, FriendliAI CEO Byung-Gon Chun explained how its cloud-based inference engine schedules and executes multiple concurrent model requests on GPUs, running “a more optimized version of continuous batching, fine grained, token level caching.”
Chun also explained larger enterprises handling a lot of traffic need to reduce the number of running GPUs, while smaller firms rely more on elasticity.
“It's more like whether they can use it if they want to use it instead of running their model all the time," Chun said. "We really focus on how to run those models very efficiently, and that's related to speed, latency, and throughput, meaning how many GPUs you need to serve many, many clients coming in. And that's really critical, and that's also connected with scalability and reliability.”
Inference Sense builds on the elastic nature of the mainline platform, going into a sleep state when not used until a new request comes in.
“This kind of interaction is just automatic from our cloud,” Chun said.
Founded in 2021 in South Korea, the Redwood City, California-headquartered FriendliAI includes Nvidia and some of South Korea’s leading telecom players in its ecosystem.
September saw it raise $20 million in a seed extension round supported by the likes of Capstone Partners, Sierra Ventures, and Korea Development Bank. FriendliAI expects its U.S. base to be 60-strong by the end of 2026, up from its current total of 45.
"Since our $20 million funding round last summer, we've been scaling our go-to-market (GTM) in the U.S. market. So we started to build GTM strongly, and the team is expanding quite quickly," Chun told this title. "With the GTM, we are hoping to help more companies in the U.S., starting from large scale, AI-native startups and also enterprise companies that use AI models, in particular, open source models."
Comments