Amazon SageMaker AI Batch Transform now supports G6e instances
Amazon SageMaker AI Batch Transform now supports EC2 G6e instances, powered by NVIDIA L40S GPUs and AMD EPYC processors, for offline inference workloads such as large language and diffusion models.
Amazon SageMaker AI Batch Transform has expanded to include support for Amazon EC2 G6e instances, enabling users to run predictions on large datasets stored in Amazon S3 without requiring a persistent inference endpoint. The G6e instances feature up to eight NVIDIA L40S Tensor Core GPUs with 48 GB of memory per GPU, alongside third-generation AMD EPYC processors, delivering enhanced performance for GPU-intensive tasks. This update allows customers to leverage G6e instances for offline inference workloads, including large language models and diffusion models used for generating images, video, and audio. The new capability is accessible through AWS SDKs, AWS CLI, or the CreateTransformJob API when selecting a supported ml.g6e instance type.
The addition of G6e instances to Batch Transform addresses the need for efficient processing of large-scale AI workloads that do not require real-time inference, offering a cost-effective solution for batch prediction tasks. By integrating these instances, users can optimize performance for computationally demanding models while maintaining flexibility in deployment. The feature is currently available in select AWS regions, including US East (N. Virginia), US East (Ohio), US West (Oregon), Asia Pacific (Mumbai), and Asia Pacific (Hyderabad).
To implement G6e instances in Batch Transform, users can initiate jobs through familiar AWS interfaces such as the AWS Management Console, AWS SDKs, or command-line tools. The process involves specifying the ml.g6e instance type during job configuration, ensuring compatibility with existing workflows. This integration simplifies the deployment of GPU-intensive models for batch inference, reducing the complexity of managing separate inference endpoints.
Amazon has provided documentation and pricing details for the new instance support, allowing customers to evaluate costs and technical specifications before adoption. Additional resources, including the Amazon SageMaker AI product page and Batch Transform documentation, offer guidance on getting started with G6e instances. Pricing information for these instances is available on the AWS pricing page.