By Interestana AI Editorial — AI-drafted, human-overseen. How we report
Google DeepMind Ships Three Physical AI Models For Robot Control
Google DeepMind has released Gemini Robotics 2, an intelligence layer for its next generation of robots, expanding capabilities beyond table-top manipulation to encompass whole-body control, five-finger dexterity, and multi-robot teamwork. This new suite is delivered as three distinct models, each with different access tiers, aiming to overcome the limitations of current robots which are often pre-programmed for narrow tasks, struggle to adapt to unpredictable environments, and exhibit poor skill transfer between different robotic bodies. Gemini Robotics 2 directly targets these three limitations simultaneously.
The three models comprising Gemini Robotics 2 are a Vision-Language-Action (VLA) model, an Embodied Reasoning (ER) Vision-Language Model (VLM), and an on-device VLA. The VLA model, named Gemini Robotics 2, translates vision and language inputs into motor control commands, enabling it to drive full humanoids from their feet to fingertips, as well as other bi-arm robots. It also facilitates dexterous manipulation, supporting both multi-finger hands and parallel grippers. The second model, Gemini Robotics ER 2, functions as the high-level "brain" for embodied reasoning. This vision-language model is designed for human communication, understanding the physical world, and planning multi-step tasks that can last for several minutes. According to its model card, ER 2 is built upon Gemini 3.5 Flash, accepting interleaved text, image, video, and audio inputs with a substantial context window of up to 128,000 tokens, and generating text outputs up to 64,000 tokens.
The third component is Gemini Robotics On-Device 2, an efficient VLA model optimized for local execution directly on the robot. Its model card indicates that it is based on Gemini Robotics 1.5 technology and Google's on-device Gemma models, processing inputs that include text, images, and robot proprioception data in numerical form. One specific checkpoint drives the Apollo 2 robot, equipped with two different types of hands, alongside a Franka Duo gripper, demonstrating the system's versatility across hardware configurations. While Gemini Robotics ER 2 is available as a public preview, the VLA and on-device models remain gated, indicating a phased rollout or specific partnership requirements for access.
Multi-finger dexterity remains an area of ongoing development, with performance metrics for this capability ranging from 32% to 92%. To address safety considerations in AI robotics, Google DeepMind has also introduced ASIMOV-Agentic, a new safety benchmark that is publicly available on Hugging Face under a CC-BY-4.0 license. This release signifies a significant step towards more adaptable, intelligent, and collaborative robotic systems capable of operating in complex and dynamic real-world environments.
Original source — read the full reporting at the publisher:
Read on MarkTechPostGet the weekly AI digest
AI news + new model releases, weekly. Drafted by our agents, reviewed by humans.