The new Amazon Elastic Compute Cloud (Amazon EC2) Trn2 instances and Trn2 UltraServers provide advanced computing capabilities for machine learning (ML) training and inference. These instances are powered by the second generation of AWS Trainium chips, known as AWS Trainium2, enhancing performance significantly over earlier models.
Trn2 instances are designed to deliver performance that is four times faster, with four times greater memory bandwidth and three times more memory capacity than the first-generation Trn1 instances. Additionally, users can expect 30 to 40 percent better price performance compared to existing GPU-based EC2 options such as P5e and P5en.
Each Trn2 instance includes 16 Trainium2 chips, offering 192 virtual CPUs (vCPUs), 2 TiB of memory, and 3.2 Tbps of Elastic Fabric Adapter v3 network bandwidth, which decreases latency by up to 50 percent when compared to earlier versions.
The Trn2 UltraServers present an entirely new compute solution, featuring 64 Trainium2 chips interconnected with a high-bandwidth, low-latency NeuronLink interconnect to optimize performance for extensive ML models.
Amazon's current infrastructure is supported by tens of thousands of Trainium chips, which are already deployed in various services. For example, on the most recent Prime Day, over 80,000 AWS Inferentia and Trainium1 chips facilitated the operation of the Rufus shopping assistant. Trainium2 chips are currently being utilized in latency-optimized configurations for top models such as Llama 3.1 and Claude 3.5 Haiku through Amazon Bedrock.
The Trn2 instances and UltraServers will be integrated in EC2 UltraClusters, enabling expansive distributed training across numerous Trainium chips on a highly efficient, non-blocking network optimized for petabit-level data transfers.
Trn2 instances are available for production in the US East (Ohio) AWS Region. Reservations can be made through Amazon EC2 Capacity Blocks for ML, allowing up to 64 instances for a maximum duration of six months and providing opportunities for instant start times and extensions as needed.
To get started with Trn2 instances, developers can utilize preconfigured AWS Deep Learning AMIs that include popular frameworks and tools such as PyTorch and JAX. Applications built with the AWS Neuron SDK can be recompiled for use on Trn2 instances, offering integration with essential libraries for distributed training and inference.
These enhancements signify substantial advancements in compute power available for machine learning and artificial intelligence workloads, streamlining the process for developers and businesses.
Comments