Alibaba Cloud unveiled a host of AI-centric offerings at its annual flagship conference, including its latest network architecture for training and inference.
Showcased at Apsara 2025, HPN8.0 is designed for AI models, with the vendor claiming it enables “seamless model training, inference, and reinforcement learning” across computational workloads.
The network architecture is capable of supporting ultra-large-scale deployments and delivers 800 Gb/s network throughput, double the capacity of its previous generation.
Alibaba Cloud provided a glimpse into the workings of HPN in a paper published in July 2024. While details on this latest version were sparse, the original iteration is Ethernet-based and features a dual-plane, two-tier design that enables interconnection of up to 15,000 GPUs within a single “pod.”
It’s designed as an alternative to traditional ECMP (equal-cost multi-path) routing. Alibaba Cloud engineers previously claimed that the prior approach led to “multiple performance issues” when handling large language model (LLM) training traffic patterns. Instead, HPN is designed to route traffic more predictably during training workloads.
HPN8.0 was among a slew of AI-focused updates at Alibaba Cloud’s event, with the vendor doubling down on upgrades of infrastructure offerings.
Eddie Wu, chairman and CEO of Alibaba Cloud Intelligence, said the updates helped to position the vendor as a “full-stack AI service provider, dedicated to delivering robust computing with maximized efficiency for training and deploying large AI models on the cloud.”
Wu also revealed Alibaba Cloud plans to invest some $53.2 billion in building out its AI and cloud infrastructure over the next three years.
The HPN8.0 reveal comes after we got a closer look at Alibaba Cloud’s software-based recovery service, ZooRoute, which it claims can “instantly” reroute network traffic during outages.
Container updates, security automation, and Vendor Buckets
Alongside the network architecture, Alibaba Cloud said it has upgraded its Container Compute Services (ACS) to enhance auto-scaling capabilities.
The offering now supports scaling of up to 15,000 pods per minute, with the vendor claiming it can handle “massive” AI agent requests. It also features improved isolation capabilities, meaning data leaks from one AI agent won’t then affect others.
Alibaba Cloud was recently named among the leadership category in Gartner’s container management rankings.
Other updates at the Apsara event saw the vendor show off enhancements to its Object Storage Service (OSS). Chief among them was the “Vendor Bucket” – an AI-powered feature designed to provide large-scale vector data storage and retrieval.
The offering unifies raw and vector data management in OSS, accessible via standard APIs to simplify retrieval-augmented generation (RAG), where LLMs are given access to specific workflows or data to improve its knowledge base.
Alibaba Cloud said Vendor Bucket provides enterprise customers with cost reductions during AI development, allowing them to manage both raw and vector data in one place.
Agentic AI features have also been added to the vendor’s Cloud Threat Detection Response (CTDR) solution, providing improved detection and response capabilities for security teams.
Some five AI agents, all powered by its Qwen line of LLMs, automate security operations such as alert assessment and reporting for end-to-end threat management.
Alibaba Cloud claimed the AI agent-based security features increased its automated incident investigation success rate from 59% to 74%, while handling 70% automated response actions without human intervention.
The vendor’s PolarDB database also received an upgrade, with improvements aimed at boosting data and AI workloads. Among the upgrades introduced was hardware innovation powered by the vendor’s Compute Express Link (CXL) technology, a compute-memory interconnect designed to reduce latency by 72.3% while boosting memory scalability.
Comments