Guozhen AIGlobal AI field notes and model intelligence

Realtime AI News

DeepMind Releases Gemini Robotics 2, Giving Robots a Gemini Brain

Google DeepMind has released Gemini Robotics 2, a family of AI models in which Gemini Robotics ER 2 acts as a cognitive brain for one or more robots, enabling planning and collaboration. The release moves embodied AI from end-to-end reinforcement learning toward a decoupled vision-language architecture, with ER 2 now available to developers via the Gemini API.

Published
DeepMind发布Gemini Robotics 2:给机器人装上Gemini大脑
Image source: deepmind.google

Google DeepMind has released Gemini Robotics 2, a family of AI models for robotics in which Gemini Robotics ER 2 acts as the "brain" for one or more robots, enabling them to plan and cooperate. According to i-programmer.info, Gemini Robotics 2 provides the intelligence layer powering the next generation of truly adaptable robots, described as a major advance that unlocks intelligent whole-body control, advanced dexterity, and multi-robot collaboration. The release is the latest step in DeepMind's nearly decade-long robotics push, which progressed from early reinforcement learning experiments in simulation to its Robotic Transformer (RT-1 and RT-2) research. Google launched the first Gemini Robotics family on top of Gemini 2.0 in March 2025, turning visual camera streams directly into robotic action, and has since partnered with third-party manufacturers including Boston Dynamics to bring Gemini's multimodal intelligence to Atlas. The most notable change is architectural. Gemini Robotics 2 moves away from relying solely on end-to-end reinforcement learning toward a decoupled, two-tier architecture driven by vision-language foundation models: a high-level cognitive layer processes streaming visual and audio input, handles open-ended language requests, calls external digital tools, and devises multi-step strategies, while a low-level execution layer of vision-language-action (VLA) models or RL-trained motor policies handles real-time spatial positioning, balance, and fine physical manipulation. The family ships in three components. Gemini Robotics ER 2, built on Gemini 3.5 Flash, serves as the cognitive brain, processing long-horizon video and audio, tracking task progress, and coordinating multi-robot teams; it is now publicly available to developers via the Gemini API and Google AI Studio, with private preview on the Gemini Enterprise Agent Platform. Gemini Robotics 2 is the primary cloud-based execution model, providing whole-body humanoid control across legs, torso, arms, and 22-DoF five-fingered hands on platforms such as Apptronik's Apollo 2. Gemini Robotics On-Device 2 runs locally on robot microprocessors without network latency and can adapt to an unfamiliar robot body in just a few hours using fewer than 200 demonstration passes. The significance is a clearer roadmap for embodied AI. Where reinforcement learning teaches a robot how to move its joints and balance its weight, vision-language foundation models decide what the robot should be doing, in what sequence, and why. For high-level cognitive tasks such as "find the spare part in the back cupboard, ask a coworker for help if it's heavy, and bring it here," designing a reward function for pure RL is virtually impossible, and the layered architecture fills that gap. Watch next: how developers adopt Gemini Robotics ER 2 through the API, how quickly the On-Device model adapts to real robots, and whether multi-robot coordination lands first in settings like factories.

Why it matters

The decoupled brain-and-body architecture gives robots a path to open-ended, long-horizon tasks that pure reinforcement learning cannot handle, and could accelerate real-world deployment of humanoid robots.

Google DeepMindGemini RoboticsRobotics
Back to realtime news

Nearby Updates

All