Google DeepMind has unveiled a new artificial intelligence model designed to act as a high-level controller for robots, marking another step towards more autonomous machines capable of carrying out complex industrial tasks with minimal human intervention.
Announced on Thursday, Gemini Robotics ER 2 is designed to sit above a robot’s existing control systems, allowing it to understand spoken instructions, interpret live video, plan multi-step workflows and coordinate multiple robots working together.
“Think of Gemini Robotics ER 2 as a high-level brain for robots,” the company said. “It allows robots to chat with humans, understand the physical world, and plan multi-step tasks. It then hands off motor execution to any given lower level vision-language-action (VLA) model.”
Unlike conventional industrial robots, which are typically programmed for fixed sequences of repetitive operations, Google said Gemini Robotics ER 2 is designed to monitor its own progress through continuous video feeds, adapt when something goes wrong and determine when a task has been completed before moving on to the next stage.
The company said the system represents a “step change” in powering robots with video understanding, task orchestration and multi-robot collaboration, enabling machines to become “more helpful in the physical world”.
The model also includes support for multiple robots working together. Google said the model allows different types of robots to communicate using a shared understanding of a task, enabling them to divide work between themselves.
One demonstration showed Boston Dynamics’ Spot robot responding to a spoken request by navigating through a building, locating a bag of popcorn and returning it to the user. Google said Gemini Robotics ER 2 orchestrated Spot’s navigation and manipulator APIs to complete the task.
Google said the model can also access external tools, including Google Search and user-defined functions, allowing robots to retrieve information as they carry out work.
A key challenge in robotics is determining whether a task has actually been completed successfully. Google said Gemini Robotics ER 2 improves robots’ ability to monitor progress throughout an operation, allowing them to adjust their actions, retry failed steps and verify that a task has been completed before proceeding.
“By watching continuous video feeds, robots can now track their own progress, adapt if something goes wrong, and know exactly when to move on to the next step,” the company said.
The company also highlighted improvements in safety, saying the latest model performs better at following safety instructions and recognising nearby people. According to Google, Gemini Robotics ER 2 can halt a humanoid robot when someone enters its working area before resuming operations once the space is clear.
The launch reflects growing competition to develop general-purpose AI systems for robotics rather than software tailored to individual machines. Instead of programming each robot separately, companies are increasingly seeking foundation models that can reason across different hardware platforms and adapt to new tasks through natural language instructions.
For manufacturers, the approach could eventually reduce the time needed to deploy and reconfigure robotic systems, particularly in environments where production lines change frequently or multiple robot types need to work together.
Google said Gemini Robotics ER 2 is now available to developers through the Gemini API and Google AI Studio, with private preview access also available through the Gemini Enterprise Agent Platform.