Remember when Microsoft and Intel were the two biggest names in IT development 25 years ago? While those two are still major players, there is no dynamic IT duo in the world right now larger than Amazon Web Services (AWS) and Nvidia.
Cloud-native developers working with the two behemoth IT companies now have more options to deploy their work. AWS and Nvidia announced this week at re:Invent 2024 that Nvidia NIM microservices are now available on AWS, supporting optimized inference and lower latency for generative artificial intelligence application deployments..
Specifically, AWS has integrated Nvidia NIM microservices across its main AI menu, which includes Amazon Bedrock Marketplace and Amazon SageMaker JumpStart. This integration enables developers to deploy optimized inference for often-used AI models, ensuring faster processing and reduced latency for building generative AI applications.
Nvidia NIM microservices are pre-built containers designed for secure and reliable deployment of AI models across various environments. Using inference engines such as Nvidia Triton Inference Server and TensorRT, these microservices support a wide range of models, from open source to custom-built, AWS CEO Matt Garman said.
Developers can deploy NIM microservices on AWS services such as Amazon EC2, Amazon EKS, and Amazon SageMaker to provide flexibility and scalability for AI workloads.
Extensive model support
A catalog of more than 100 NIM microservices, including models such as Meta's Llama 3 and Mistral AI's Mistral, is now available for preview. These microservices are optimized for Nvidia-accelerated computing instances on AWS, ensuring optimal performance.
This type of software development isn't for the faint of heart, and anything that can be done to save time and repetitive actions is always welcome. Key benefits of all this for developers, according to AWS, include: Other AWS-Nvidia highlights from re:Invent 2024 include:
- Accelerated AI inference: Reduced latency and improved performance for generative AI applications
- Simplified deployment: Easy access to pre-built, optimized inference solutions
- Broad model support: Compatibility with a wide range of AI models
- Scalability and flexibility: Deployment across various AWS services
- Enhanced developer productivity: Streamlined workflows for AI development
- Nvidia DGX Cloud on AWS: Nvidia announced DGX Cloud on AWS, providing enterprises with a high-performance solution for AI model training and inference, plus direct access to Nvidia expertise. Early adopter Leonardo.ai is already using the platform to accelerate AI initiatives
- AWS liquid-cooled data centers with Nvidia Blackwell; New P6 Instances: AWS's next-generation data centers will feature liquid cooling technology for Blackwell GPUs. Also, new AWS P6 instances with Nvidia Blackwell will be the foundation for Amazon EC2, DGX Cloud, and Project Ceiba
- Nvidia expands robotics simulation on AWS with Isaac Sim: Nvidia announced Isaac Sim availability on Amazon EC2 G6e instances with L40S GPUs, enabling companies like Vention and Field AI to simulate and test AI-driven robots in physically-based virtual environments
- Nvidia CUDA-Q on AWS Braket for quantum computing: Nvidia has integrated CUDA-Q with AWS Braket, simplifying hybrid quantum-classical application development across multiple quantum processors
- Nvidia brings edge AI to AWS IoT with IGX and Jetson Integration: Nvidia announced the integration of IGX Orin and Jetson Orin platforms with AWS IoT Greengrass, streamlining AI model deployment and device management at the edge
Comments