South Korean operator SK Telecom (SKT) claimed it can solve memory supply chain issues using SK Hynix wares as it continues to solidify its AI operations following the firm's major reorganization last year.
In an exclusive media briefing hosted at this month’s Mobile World Congress (MWC), the firm broke down its progress since splitting into two parts to concentrate its efforts in the AI space, with one arm operating all things telecom and the other all things AI through the creation of new division, AI-CIC (Company-in-Company).
Min-young Jeong, head of AI data center solutions, explained that firms need to carry out their own inferencing due to using a variety of models, including open-source options, and “won’t find it easy” – hence SKT touting an all-in-one stack on the MWC show floor. The AI DC head also stressed that the firm’s new operating model through AI-CIC is not a one-stop operation, with SKT instead taking a phased approach while keeping its colocation business model.
SKT leverages sister brand SK Hynix, with the firm touting Hynix's Gaia generative AI platform optimized for semiconductor operations. Gaia is included in the SKT stack and tracks metrics such as GPU utilization, high-bandwidth memory usage, temperature, power, and PCIe/network traffic through gauges and time-series graphs.
“With the in-rack domain, we focus on resolving memory bottlenecks in AI workloads, leveraging SK Hynix’s world leading high-bandwidth memory (HBM) technology,” Jae-shin Lee, head of AI business development for SKT, explained, touching on the so-called memory wall issue where a performance gap is growing between fast processors and relatively slow memory, impacting system performance as processors spend significant time waiting for data.
With Jeong revealing that disaggregation was a “key word” for the company in the year ahead, Lee referred to SKT's disaggregated inference platform, claimed to maximize performance and lower costs while optimizing resources allocation for inference tasks. Specifically, the platform decouples and decodes inference stages for optimized resource allocation, separating the two distinct phases of large language model (LLM) processing: prefill, or input processing, and decode, or token generation, as split across dedicated, specialized hardware resources. This approach can optimize performance by separating compute-bound prefill from memory-bound decoding, allowing independent scaling, reduced latency, and improved GPU utilization during the inference process.
SKT also pointed to their platform leveraging specialized hardware types for improved token-per-dollar efficiency with an AI computing rack made up of "open composable fabric" for high-density AI compute and horizontal memory scaling, alongside integrated infrastructure management, including power solutions and AI-native DCIM.
Hangul and Haein
On the Barcelona show floor, SDxCentral also saw SKT’s work on the A.X K1 LLM hyperscale AI model. A.X K1 was developed under a government-backed Sovereign AI Foundation Model project, and is claimed to be the first South Korean hyperscale model with 519 billion parameters.
The SKT consortium plans to bolster nationwide AI accessibility by offering A.X K1 through A. (A-DoT), its AI assistant service deployed across both consumer and business platforms. For the former, the consortium aims to build an "AI for Everyone" framework, enabling the South Korean public to easily access AI via phone calls, text messages, the web, and apps. One of its developers also revealed that it features a tokenizer for Korean script that’s 33% more efficient than using ChatGPT in Korean language mode.
Furthering SKT’s sovereign angle is its Haein GPU cluster, which launched in August as is to be the backbone of South Korea’s designated sovereign AI infrastructure through providing GPU-as-a-service (GPUaaS). The cluster, claimed to be South Korea’s largest, is powered by more than one thousand Nvidia B200 GPUs, with flexible GPU allocation driven by Petasus AI Cloud, a platform that virtualizes GPU clusters and the interconnect fabric, such as NVLink, InfiniBand, and remote direct memory access over converged Ethernet version 2 (RoCEv2).
Comments