OFICIAL AWS What's New

Amazon SageMaker AI Batch Transform now supports G6e instances

What happened
Based on AWS What's New · Sep 04, 2026

Amazon SageMaker AI Batch Transform now supports EC2 G6e instances, powered by NVIDIA L40S GPUs and AMD EPYC processors, for offline inference workloads such as large language and diffusion models.

Amazon SageMaker AI Batch Transform now supports G6e instances
AWS What's New — Amazon Web Services
Key points
·
Amazon SageMaker AI now supports Amazon EC2 G6e instances for batch transform.
·
Batch Transform enables you to run predictions on datasets stored in Amazon S3 and is suited for large datasets that do not require a persistent inference endpoint.
·
Amazon EC2 G6e instances are powered by up to eight NVIDIA L40S Tensor Core GPUs with 48 GB of memory per GPU and third-generation AMD EPYC processors.
·
G6e instances deliver improved performance for GPU-intensive workloads.
Key numbers
·
Amazon SageMaker AI Batch Transform has expanded to include support for Amazon EC2 G6e instances, enabling users to run predictions on large datasets stored in Amazon S3 without requiring a persistent inference endpoint.
·
The G6e instances feature up to eight NVIDIA L40S Tensor Core GPUs with 48 GB of memory per GPU, alongside third-generation AMD EPYC processors, delivering enhanced performance for GPU-intensive tasks.
·
g6e instance type during job configuration, ensuring compatibility with existing workflows.

Amazon SageMaker AI Batch Transform has expanded to include support for Amazon EC2 G6e instances, enabling users to run predictions on large datasets stored in Amazon S3 without requiring a persistent inference endpoint. The G6e instances feature up to eight NVIDIA L40S Tensor Core GPUs with 48 GB of memory per GPU, alongside third-generation AMD EPYC processors, delivering enhanced performance for GPU-intensive tasks. This update allows customers to leverage G6e instances for offline inference workloads, including large language models and diffusion models used for generating images, video, and audio. The new capability is accessible through AWS SDKs, AWS CLI, or the CreateTransformJob API when selecting a supported ml.g6e instance type.

The addition of G6e instances to Batch Transform addresses the need for efficient processing of large-scale AI workloads that do not require real-time inference, offering a cost-effective solution for batch prediction tasks. By integrating these instances, users can optimize performance for computationally demanding models while maintaining flexibility in deployment. The feature is currently available in select AWS regions, including US East (N. Virginia), US East (Ohio), US West (Oregon), Asia Pacific (Mumbai), and Asia Pacific (Hyderabad).

To implement G6e instances in Batch Transform, users can initiate jobs through familiar AWS interfaces such as the AWS Management Console, AWS SDKs, or command-line tools. The process involves specifying the ml.g6e instance type during job configuration, ensuring compatibility with existing workflows. This integration simplifies the deployment of GPU-intensive models for batch inference, reducing the complexity of managing separate inference endpoints.

Amazon has provided documentation and pricing details for the new instance support, allowing customers to evaluate costs and technical specifications before adoption. Additional resources, including the Amazon SageMaker AI product page and Batch Transform documentation, offer guidance on getting started with G6e instances. Pricing information for these instances is available on the AWS pricing page.

Original source → Deals on Clipraptor.com →