Gemini Robotics 2 brings whole body intelligence to robots
Google DeepMind unveils Gemini Robotics 2, an AI model enabling robots to perform whole-body control, advanced dexterity, and multi-robot collaboration, with on-device operation and safety enhancements.
Video
Video available
The useful question is what changes for users, developers or buyers, and whether the announcement stays industry context or becomes something people can actually use.
Gemini Robotics 2 introduces a unified AI model that allows robots to reason through full-body movements, enabling tasks such as navigating cluttered spaces, manipulating objects, and collaborating with other robots. The system supports diverse robotic platforms, including humanoid robots like Apptronik’s Apollo 2, and can adapt to new robot designs with minimal training data. This marks a shift from pre-programmed or teleoperated robots to AI-driven systems capable of learning and adapting in real-world environments.
The model enhances physical dexterity by controlling complex end effectors, such as multi-fingered hands for delicate tasks like tying knots or sealing bags, and standard grippers for precision tasks like packing. While whole-body and gripper-based tasks show medium to high success rates, multi-finger manipulation remains challenging. The system also supports multi-robot workflows, allowing different robots to coordinate on complex tasks that exceed individual capabilities.
Gemini Robotics ER 2, the reasoning layer, processes user instructions, plans multi-step tasks, and monitors progress, enabling robots to execute sequences lasting several minutes with hundreds of decisions. It introduces improved task tracking, self-correction, and generalization to novel situations. The model is available on Google AI Studio and in private preview on the Gemini Enterprise Agent Platform, with early-access versions for on-device deployment.
Safety is a core focus, with advancements in embodied reasoning to detect human proximity, halt unsafe actions, and request human intervention when necessary. The ASIMOV-Agentic benchmark evaluates agentic safety, including refusal of unsafe actions and proactive human involvement. The update also includes an on-device vision-language-action model optimized for latency-free operation, supporting rapid adaptation to new robot embodiments with minimal examples.