Gemma-4-31B-it-assistant and Gemma-4-31B-IT-NVFP4 models now available on Amazon SageMaker JumpStart
AWS now offers Google DeepMind’s Gemma-4-31B-it-assistant and NVIDIA’s Gemma-4-31B-IT-NVFP4 models on SageMaker JumpStart for enterprise AI deployments.
Amazon SageMaker JumpStart now hosts Google DeepMind’s Gemma-4-31B-it-assistant and NVIDIA’s Gemma-4-31B-IT-NVFP4, adding two new foundation models to AWS’s catalog. The assistant-tuned Gemma-4-31B-it-assistant supports multimodal inputs including text, images, and video frame sequences, with a 256K-token context window and over 140 language support. It ranks third on the Arena AI text leaderboard among open models and features a hybrid attention mechanism with native function calling for agentic workflows.
NVIDIA’s Gemma-4-31B-IT-NVFP4 delivers the same Gemma 4 31B capabilities in a quantized 4-bit FP4 format, reducing memory usage to approximately 18.5 GB—68% smaller than the base model. The quantized variant achieves roughly 2.5x faster inference while maintaining 97–99% of the original model’s quality, making it suitable for high-throughput production deployments.
Both models are optimized for AWS infrastructure and can be deployed via SageMaker JumpStart with minimal setup, enabling customers to integrate them into existing workflows efficiently. The assistant variant is designed for reasoning, coding, and agentic tasks, while the quantized model targets cost-efficient, scalable deployments on NVIDIA GPUs.
Customers can access these models through the SageMaker JumpStart model catalog in the SageMaker console or via the SageMaker Python SDK for programmatic deployment. AWS provides documentation and guidance for deploying and using foundation models within SageMaker JumpStart.