samsung
– Giacomo Lee/SDxCentral

Inference platform FriendliAI is partnering with Samsung’s IT division to offer Nvidia GPU-based frontier AI services.

FriendliAI's core Friendli Inference will be deployed by Samsung SDS on its Samsung Cloud Platform (SCP), which runs on Nvidia DGX B300 GPU infrastructure, offering support for models such as GLM-5, MiniMax M2.5, and DeepSeek v3.2. The collaboration promises SCP’s mainly South Korea-based customers high-performance, low-latency inference, and so-called "frontier AI."

“We are pleased to collaborate with Samsung SCP to provide high-performance, high-efficiency AI inference to companies worldwide, ” FriendliAI CEO Byung-Gon Chun told South Korean publication Naver. “Through our collaboration with SCP’s cutting-edge Nvidia B300 GPU infrastructure, customers will be able to reliably utilize the latest frontier models and explore new business opportunities based on agentic AI.”

"Samsung SCP provides high-performance GPU infrastructure as infrastructure-as-a-service (IaaS) and is continuously expanding environments optimized for AI workloads," Samsung SDS VP Eun-young Kim added. "Through collaboration with Friendly AI, we will provide frontier model inference services … and offer more stable and scalable AI infrastructure to global customers."

Founded in 2021 in Seoul, South Korea, the Redwood City, California-headquartered FriendliAI includes Nvidia and some of South Korea’s leading telecom players within its ecosystem.

Last September saw it raise $20 million in a seed extension round supported by the likes of Capstone Partners, Sierra Ventures, and Korea Development Bank. FriendliAI expects its U.S. base to be 60-strong by the end of 2026, up from its current total of 45.

Most recently, the firm launched an "Adsense for GPUs" designed to tackle idle AI inferencing. Specifically, the Friendli InferenceSense tool detects idle GPU capacity in a user’s infrastructure to fill it with monetizable AI inference workloads via secured, fully-isolated containers that serve paid inference workloads, while the inference engine maximizes token throughput per GPU-hour.