HeadlinesBriefing favicon HeadlinesBriefing.com

Google Gemini Robotics 2.0 Enhances Robot Capabilities

Ars Technica •
×

Google has unveiled Gemini Robotics 2.0, a significant advancement in AI-powered robotics. This new iteration features three sub-models, with one, Gemini Robotics ER 2, now available to developers. ER 2 is a vision language model (VLM) designed to understand complex instructions and analyze its environment through live video feeds, processing video frames with nearly 60 percent accuracy and identifying key task moments with almost 90 percent accuracy. This allows robots to perform multi-step tasks with improved dexterity and to recover from failures in real-time, such as readjusting a grip without restarting the entire task.

Gemini Robotics 2.0 also introduces whole-body intelligence, enabling robots to execute actions based on instructions. This is managed by another model, Gemini Robotics 2, which generates robot actions. An offline version, Gemini Robotics On-Device 2, also exists. Google is testing these models on hardware from Boston Dynamics and Apptronik, demonstrating improved accuracy and efficiency. Even the on-device version can adapt to new robot designs with limited data, around 200 examples.

Safety remains a core focus, with Gemini Robotics 2 incorporating traditional physical safety measures and robust AI frameworks. A new safety benchmark, ASIMOV-Agentic, evaluates models on factors like refusing unsafe actions and requesting human assistance. Gemini Robotics ER 2 is highlighted as the safest model yet, capable of understanding safety protocols and halting actions when humans are in close proximity.