Amazon Web Services has launched a slew of updates to help customers power their AI workloads.
Chief among them was the news that AWS EC2 customers can finally get their hands on Nvidia’s next-gen Blackwell hardware to power their AI workloads.
The cloud computing giant also rolled out enhancements to SageMaker, its platform that helps users build and train AI models, with new observability features, and streamlined deployment workflows.
AWS unveils Blackwell-powered instances for AI training and inference
To power customer training and inference workloads, AWS unveiled two new system configurations: the P6-B200 and P6e-GB200, both of which are now generally available.
The P6e-GB200 - which AWS bills as its most powerful EC2 instance ever – houses 36 Grace CPUs, and 72 Blackwell GPUs, connected using Nvidia’s proprietary fifth-generation NVLink system.
AWS claims its new instances can provide users with up to 360 petaflops of FP8 (8-bit floating point) compute and 13.4 TB of total high bandwidth memory (HBM3e), with the hyperscaler touting it as a rack-scale supercomputer capable of supporting the latest AI foundation models.
The hyperscaler opted for its in-house networking stack, the Elastic Fabric Adapter (EFAv4), powered by its custom Nitro controllers, which optimizes network packet processing and reduces latency for high-throughput applications.
AWS claims its networking picks mean the P6e-GB200 can offer up to 28.8 Tbps of networking bandwidth per server. The interconnect offering means users can effectively scale workloads to encompass tens of thousands of Blackwell GPUs.
“P6e-GB200 UltraServers are ideal for the most compute and memory-intensive AI workloads, such as training and inference of frontier models, including mixture of experts models and reasoning models, at the trillion-parameter scale,” an AWS blog post reads.
The air-cooled P6-B200s, meanwhile, are suited for less sizable training and inference applications, though still pack a punch, with the hyperscaler touting them offering up to 2x performance compared to the prior generation Hopper-based P5en instances for AI training and inference.
Customers can use the SageMaker HyperPod solution (more below) to manage their choice of instance, handling provisioning and management of their cluster via built-in dashboards. It also automatically replaces faulty instances in the same NVLink domain, which can help to keep workloads running and reduce data drops.
David Brown, VP of AWS compute and machine learning services, wrote in a post: “This launch announcement is an important milestone, and it’s just the beginning. As AI capabilities evolve rapidly, you need infrastructure built not just for today’s demands but for all the possibilities that lie ahead.
“With innovations across compute, networking, operations, and managed services, P6e-GB200 UltraServers and P6-B200 instances are ready to enable these possibilities.
The launch of the new EC2 instances sees AWS become the latest cloud provider to get its hands on Blackwell. Nvidia's flagship architecture is only now making its way out to customers after a sizable delay put a halt to production. Earlier this week, CoreWeave became the first cloud vendor to get its hands on the next iteration of Blackwell, the highly anticipated Nvidia GB300 NVL72.
The increasing variety of GPUs powering AWS's EC2 instances now includes the GB200 NVL72. The G family, utilized for graphics and ML inference, already incorporates hardware from both Nvidia and AMD.
GB200 NVL72 for its EC2 platform comes as it’s also using the rack-scale solution to power Project Ceiba, its attempt to build the world's largest cloud AI supercomputer. Ceiba had initially featured Grace Hopper Superchips before receiving a Blackwell-based upgrade last March.
The launch of its P6 offerings comes just a week after AWS unveiled new network-optimized instances. Based on the hyperscaler's custom Graviton4 processors, the EC2 C8gn instances offer up to 30 percent higher compute performance compared to the prior-gen C7gns.
SageMaker AI gets observability boost, development workflow enhancements
Alongside its infrastructure updates, AWS rolled out a series of enhancements to SageMaker AI designed to streamline AI model development workflows and reduce time-to-market for organizations building generative AI applications.
Chief among the new features is SageMaker HyperPod's observability capability, which AWS claims can reduce troubleshooting time from days to minutes.
The feature provides a unified dashboard preconfigured in Amazon Managed Grafana, with monitoring data automatically published to an Amazon Managed Service for Prometheus workspace.
The observability solution addresses a common pain point for data scientists and ML engineers who previously had to manually browse through entire cluster resources to correlate job failures with hardware issues.
Teams using the platform can now quickly filter monitoring data for specific GPUs that performed failed jobs, spot bottlenecks, and optimize compute resources through a single interface.
AWS has also simplified the path from model development to production inference with the ability to deploy Amazon SageMaker JumpStart models directly on SageMaker HyperPod.
The feature allows data scientists to run inference on models with a single click, eliminating manual infrastructure setup and reducing large model downloads from hours to minutes.
AWS claims that, having trained its own Nova foundation models on SageMaker HyperPod, the company saved months of work and achieved compute resource utilization of more than 90 percent.
Customers, including Hugging Face, Perplexity AI, and Salesforce, are also already leveraging SageMaker HyperPod for model development.
AWS also introduced remote connections, allowing developers to connect their local Visual Studio Code environments to SageMaker AI. The feature enables developers to maintain their customized local setups while accessing SageMaker AI's infrastructure, scalability, and security controls.
Additionally, AWS has launched a new command line interface (CLI) and software development kit (SDK) for SageMaker HyperPod, providing a consistent interface that simplifies infrastructure management and unifies job submission across training and inference workloads.
Ankur Mehrotra, GM for Amazon SageMaker AI, wrote: “Since launching in 2017, SageMaker AI has transformed how organizations approach AI model development by reducing complexity while maximizing performance.
“Since then, we’ve continued to relentlessly innovate, adding more than 420 new capabilities since launch to give customers the best tools to build, train, and deploy AI models quickly and efficiently.”
Comments