
Google launches Gemini Robotics ER 2 for video-aware, multi-robot physical AI
Google's Gemini Robotics ER 2 adds video progress tracking, task orchestration and multi-robot collaboration for physical AI developers.
Google has introduced Gemini Robotics ER 2, a new embodied reasoning model aimed at giving robots a higher-level system for planning tasks, interpreting live video and coordinating with other machines. The company says the model is available now through the Gemini API and Google AI Studio, with private preview access through the Gemini Enterprise Agent Platform.
The July 30 announcement positions ER 2 as a control layer above lower-level robot action systems. Instead of replacing a robot's movement controller, the model is designed to read multimodal input, plan multi-step work, call tools or user-defined functions, and hand execution to robotics APIs or vision-language-action models. Google says this lets a robot keep reasoning about the next step while a task is already underway, reducing stop-and-wait behavior that can make physical automation brittle.
Why The Update Matters
The biggest practical change is the model's use of continuous video to judge whether a job is progressing correctly. Google says ER 2 can classify task progress in five bands and reached 57.4 percent accuracy on its progress-classification tests. For moment-finding, the task of identifying the exact frame where an event happens, Google reports 91.3 percent accuracy and a 0.96-second mean absolute distance. Those claims matter because a robot pouring, tightening, carrying or handing off an object needs to know when to continue, stop or ask for help.
Google also highlighted multi-robot collaboration. In its examples, different machines can share a semantic understanding of a workspace and divide work that one robot could not complete alone. The announcement names demonstrations involving Boston Dynamics Spot and partner robots including Apptronik's Apollo 2 and Franka's F3 Duo, but the broader pitch is for developers to connect ER 2 to their own control interfaces.
Availability And Limits
For developers, the near-term value is experimentation rather than a finished consumer robot. Google says examples are available on GitHub, while enterprise access is still in private preview. The company is also making safety a major part of the claim: ER 2 is described as improving safety instruction following and human-proximity behavior, including halting a humanoid robot when a person is nearby and resuming only after the area is clear.
Those are still company-reported benchmarks, so real-world performance will depend on hardware, sensors, integration quality and the tasks developers attempt. Even so, ER 2 is a meaningful sign that frontier AI labs are moving from chat and software agents toward systems that can monitor and orchestrate physical work in real time.
Sources
Cover photo by Freek Wolsink on Pexels, used under the Pexels License.
CyberOGZ Team






Comments (0)
Leave a Comment