Introducing Gemini Robotics ER 2
Google launched Gemini Robotics ER 2, a new model designed to serve as a high-level control system for robots, enabling real-time spatial reasoning, multi-step task planning, and multi-robot collaboration via the Gemini API and related platforms.
The useful question is what changes for users, developers or buyers, and whether the announcement stays industry context or becomes something people can actually use.
Google has introduced Gemini Robotics ER 2, a model intended to function as a central decision-making system for robots. It supports real-time spatial reasoning, multi-step task planning, and collaboration between different robots. Developers can access the model through the Gemini API, Google AI Studio, or the Gemini Enterprise Agent Platform to build physical AI agents. The model processes continuous video feeds to track progress, adapt to errors, and determine when to proceed to the next task step.
Gemini Robotics ER 2 is positioned as an upgrade over its predecessor, ER 1.6, with enhanced capabilities for video-based progress tracking and self-correction. It enables robots to perform complex workflows by orchestrating lower-level vision-language-action models and APIs. The model can call external tools such as Google Search or user-defined functions, and it integrates with the Gemini Live API for low-latency task execution. A demonstration with Boston Dynamics’ Spot robot illustrates its ability to fetch objects based on natural language commands.
The model introduces improvements in task progress understanding, including progress classification and moment-finding. Progress classification assigns video frames to discrete completion stages, allowing robots to adjust actions dynamically. Moment-finding identifies precise moments for task transitions, such as stopping a pouring action. ER 2 achieves 57.4% accuracy in progress classification and 91.3% in moment-finding, with sub-second latency required for real-world robotics operations.
Gemini Robotics ER 2 also supports multi-robot collaboration, enabling diverse machines to work together using a shared semantic framework. Safety features include adherence to physical constraints and human proximity detection, with the model halting operations when humans are nearby. ER 2 outperforms ER 1.6 and other models in safety benchmarks, including Safety Instruction Following and Human Proximity. The model is available to developers for building and testing physical AI applications.