sakutto
Generative AI· Gemini Robotics 2

Gemini Robotics 2: Whole-Body Control and Multi-Robot Teamwork

GeminiGoogle DeepMindroboticsVLA
Gemini Robotics 2: Whole-Body Control and Multi-Robot Teamwork

What Gemini Robotics 2 is

Gemini Robotics 2 is a family of AI models for driving robots, announced by Google DeepMind on July 30, 2026. The company positions it as the intelligence layer powering the next generation of adaptable robots.

Robot control has long meant either pre-programming a fixed sequence or having a human teleoperate. DeepMind takes that as its starting point and names two gaps: the inability to genuinely learn or adapt to unpredictable environments, and the fact that transferring a learned skill from one robot body to another remains extremely difficult. This release goes after both.

View official source →
Most robots are pre-programmed or teleoperated for narrow, repetitive task sequences. They lack the ability to truly learn for themselves or adapt to unpredictable environments. Moreover, transferring learned skills from one robot body to another remains incredibly difficult. / Today, we are introducing Gemini Robotics 2 - the intelligence layer powering the next generation of truly adaptable robots. — From the passages on the limitations of conventional robot control and the positioning of Gemini Robotics 2

From upper body to whole body

The headline change is how much of the body it can drive. Earlier models controlled a humanoid's upper body to get table-top tasks done. Now walking, crouching and reaching connect to object manipulation as a single continuous flow.

DeepMind's example is asking Apptronik's Apollo 2 humanoid to "put the watering can into the green bin in the bottom shelf." The robot interprets the instruction, walks to the table, picks up the watering can, takes a few steps to the shelves, and places it. What is new is that walking, grasping and placing happen as one response to one instruction rather than as separate capabilities.

View official source →
While our previous models controlled the humanoid's upper-body to achieve table-top tasks, Gemini Robotics 2 expands physical AI into whole-body motions. / For example, when controlling Apptronik's Apollo 2 humanoid robot, we can ask it to "put the watering can into the green bin in the bottom shelf." Apollo processes the instruction, walks to the table, and picks up the watering can, takes a few steps to the shelves, and places it precisely in its destination. — From the passages on the expanded control range and the demonstrated task

A "VLA" turns sight and language into motion

At the core of Gemini Robotics 2 is a class of model called a VLA—vision-language-action.

It does what the name says: it takes what the cameras see plus a spoken instruction, and converts them into signals for how to move the motors. Where a text-generating model outputs the next word, a VLA outputs the next movement. It can drive a full humanoid from feet to fingertips, and also handles bi-arm robots.

View official source →
Gemini Robotics 2: Our most advanced vision-language-action model (VLA) that converts vision and language input into motor control, enabling a robot to take action. This model is capable of controlling full humanoids, from feet to fingertips, and other bi-arm robots. It also brings a new level of dexterous manipulation on both hands and grippers. — From the definition of the VLA model and the robot types it supports

The same checkpoint drives different bodies

Here is the answer to that skill-transfer problem. The results chart DeepMind published comes from controlling three different robots with one and the same model checkpoint (the saved internal state of a trained model).

The three are Apollo 2 with SharpaWave hands, Apollo 2 with Inspire hands, and a Franka Duo with a Robotiq gripper. A gripper is a device that grasps by clamping; the one used here is a standard two-finger parallel type. Two are the same humanoid with different hand hardware, and the third is not a humanoid at all. The claim is that one model covers multiple bodies instead of each body needing its own.

View official source →
Gemini Robotics 2 controlling three different embodiments, using the same model checkpoint — the Apptronik Apollo 2 robot with SharpaWave hands, the Apollo 2 robot with Inspire hands, and the Franka Duo with the Robotiq gripper — on a wide variety of whole-body and dexterous manipulation tasks. — From the description of three embodiments driven by a single checkpoint

How the three models divide the work

Gemini Robotics 2 is not one model but three with different jobs. Getting this wrong leads to wrong conclusions about what you can try today, so it is worth laying out first. In the table, VLA is the vision-language-action model described above and VLM is a vision language model.

The three models that make up Gemini Robotics 2 (per the official announcement)

ModelTypeWhat it handles
Gemini Robotics 2VLAConverts vision and language into motor control and actually moves the body
Gemini Robotics ER 2VLMTalks with humans, understands the physical world, plans multi-minute sequences
Gemini Robotics On-Device 2VLALightweight version running on the robot itself; handles adaptation to new bodies
View official source →
Gemini Robotics ER 2: Our most capable embodied reasoning (ER) model. It is a vision language model (VLM) that acts as our agent, enabling robots to communicate with humans, understand the physical world and plan multi-step tasks lasting several minutes. We are also introducing the ability for robots to work together as a team. / Gemini Robotics On-Device 2: Our most efficient vision-language-action model (VLA) optimized to run locally on robotic devices. This model can now achieve fast adaptation to completely new robot embodiments with a few hours of data. — From the descriptions of the type and scope of ER 2 and On-Device 2

ER 2 is the thinking half

ER stands for embodied reasoning—reasoning while having a body. ER 2 is a VLM (vision language model) that acts as the higher-level brain for the robot.

It observes the room, assembles the steps a task needs, coordinates with the VLA to carry them out, and tracks progress to completion. DeepMind says this arrangement lets robots self-correct when a step fails and generalize to situations and goals they have not seen. The horizon is several minutes and hundreds of decisions. Understanding when tasks begin and end, and pinpointing the moment key events occur, are also listed as improvements in this update.

View official source →
It observes the room, reasons about the steps needed to complete the task, coordinates with the VLA to carry out the actions, and tracks progress until the task is done. This setup allows robots to execute complex multi-step tasks, self-correct if a step fails, and generalize to novel situations and goals. / In this update, we are enabling robots to more reliably execute longer task sequences, lasting several minutes and involving hundreds of decisions. Gemini Robotics ER 2 now understands when tasks begin and end, and can pinpoint the moment key events occur, marking a step change in progress understanding. — From the role of the ER model, the task lengths it handles, and its understanding of task start and end

Different kinds of robots can work together

Multi-robot collaboration arrives alongside. Different types of robots can now communicate and split up work a single robot could not complete.

DeepMind's example is tidying a cluttered room, where a robot "can even team up with other robots to finish the job faster." The framing is not about making one humanoid more capable—it is about being able to add units when there are not enough hands.

View official source →
It can even team up with other robots to finish the job faster. / Furthermore, we are introducing multi-robot collaboration. This enables different types of robots to communicate and work together to solve complex workflows a single robot could not do alone. — From the statements on teaming up with other robots and on introducing multi-robot collaboration

On-Device 2 exists for places with no connection

The third model, On-Device 2, is a lightweight VLA optimized to run on the robot itself, aimed at environments where network latency or internet access cannot be relied on.

Its standout role is adaptation to new bodies. Using the "motion transfer" techniques inherited from Gemini Robotics 1.5, DeepMind says it adapts to new bi-arm robots with a few hours of adaptation time and typically under 200 examples. This is said to hold even for bodies with drastically different shapes, sensors and degrees of freedom (how many directions the joints can move in), demonstrated across the Dexmate, SO101 and Trossen platforms.

Specifications like these are worth checking against the source text, which keeps you from missing the qualifiers. Converting the official blog into a readable, searchable form makes it easier to find the passages later.

Free ToolURL to Markdown ConverterConvert any public web page URL to Markdown. Preserves headings, tables, lists, and links — perfect for LLM and RAG preprocessing, research notes, and archiving web articles.Try it now →

View official source →
This model is natively multi-embodiment and inherits our advanced "motion transfer" techniques from Gemini Robotics 1.5. We can now adapt to new bi-arm robot embodiments with just a few hours of adaptation time, typically with less than 200 examples. This works even with new embodiments with drastically different shapes, sensors and degrees of freedom, as shown below with a diverse set of tasks being performed by the Dexmate, SO101, and Trossen platforms. — From the on-device model's adaptation speed and the platforms it was demonstrated on

The limits DeepMind states outright

This is the part of the announcement most worth reading. Alongside the results, the company is explicit about what does not work yet.

Multi-finger dexterity is still hard

The caption on the results chart states that while whole-body and gripper-based dexterous tasks reach a medium to high success rate, multi-finger dexterous manipulation remains challenging. The body of the announcement describes reaching the point of tying knots and sealing a ziplock bag with the five-fingered, 22-degree-of-freedom SharpaWave hand—but by DeepMind's own assessment, that is not yet at a reliably repeatable level.

Whether a robot can do fine work with its fingers is what separates useful from not in homes and workplaces. Between "the demo video shows a knot being tied" and "you can hand that job over" sits exactly this gap in success rate.

View official source →
While Gemini Robotics 2 achieves a medium to high success rate for whole-body and gripper-based dexterous tasks, the multi-finger dexterous manipulation remains challenging. / The model can now control the five-fingered, 22 degree-of-freedom SharpaWave hand on the Apollo 2 robot to complete delicate actions like tying knots or sealing a ziplock bag. — From the passages on the multi-finger challenge and the SharpaWave hand specification

Movement speed has room to improve

Immediately after the watering-can demonstration comes a caveat that movement speed still needs to advance. DeepMind frames the result as an important step toward real-world tasks requiring whole-body coordination, while stating that speed remains an open problem.

Being able to decide and act autonomously is a separate question from acting at human speed. For anyone evaluating deployment, the question is not only whether the task completes but whether it completes in a practical amount of time.

View official source →
While our robots have more to advance in movement speed, this is an important step towards the skills needed to complete more complex, real-world tasks that require whole-body coordination. — From the caveat regarding movement speed

Safety is built as "the reasoning model stops the acting model"

On safety, a new benchmark called ASIMOV-Agentic arrives. What it measures is whether the reasoning ER model can refuse unsafe tool calls coming from the acting VLA. A tool call here means the AI triggering an external capability itself—in this case, robot motion. It also evaluates whether the agent can predict that a task is impossible, and whether it proactively asks for human intervention when uncertain.

The design puts safety in two layers, with the thinking side supervising the acting side. DeepMind calls ER 2 its safest robotics model to date on safety-constraint-following and human-proximity benchmarks, and notes that detecting an approaching human and bringing the robot to a safe stop is a key requirement in collaborative safety standards.

View official source →
We're introducing ASIMOV-Agentic, a new benchmark for agentic safety orchestration and uncertainty resolution. For example, it measures the embodied reasoning agent's ability to refuse unsafe tool calls from a VLA. It also measures the agent's ability to predict whether a task is possible and to proactively request human intervention when uncertain. / Additionally, with enhanced embodied reasoning, Gemini Robotics ER 2 is our safest robotics model to date in safety constraint following and human proximity benchmarks. It can better detect when humans are nearby, trigger safety tool calls and bring the robot to a safe stop if someone approaches too closely. This is a key requirement in collaborative safety standards. — From what the ASIMOV-Agentic benchmark measures and the human-proximity safety results

How much of it you can actually try

"Announced" and "available" are different things, and the three models differ in availability.

Only ER 2 is generally reachable

The reasoning model, Gemini Robotics ER 2, is available in Google AI Studio and in private preview on Gemini Enterprise Agent Platform. If you have tried a Gemini model in AI Studio, it is the same entry point.

The VLA model that actually drives motors and the on-device model are limited to early-access partners. Having a robot on hand does not mean you can run Gemini Robotics 2 on it today. Those who want to try are pointed to the Trusted Tester Program.

Free ToolURL to Markdown ConverterConvert any public web page URL to Markdown. Preserves headings, tables, lists, and links — perfect for LLM and RAG preprocessing, research notes, and archiving web articles.Try it now →

View official source →
Gemini Robotics ER 2, our reasoning model, is now available on Google AI Studio and in private preview on Gemini Enterprise Agent Platform. Our VLA and On-Device models are available to early-access partners. — From the statement of availability for the three models

Read it as part of general-purpose models reaching into robots

Handing robot control to a general-purpose AI model rather than a task-specific one is not unique to Google DeepMind. Anthropic published Project Fetch, an experiment in which a general-purpose model with no robotics training operated an off-the-shelf robot dog.

The emphases differ. Project Fetch measures how far you get with no special training, while Gemini Robotics 2 is a family built for robot control from the start. But both descend from the same Gemini foundation lineage, and both show language-model progress flowing straight into the physical world. DeepMind itself calls this a milestone on the path toward AGI (artificial general intelligence—AI that handles tasks across the breadth a human does) in the physical world.

View official source →
Gemini Robotics 2 marks an important milestone on the path toward solving AGI in the physical world. Unlocking the true potential of robotics requires moving past single-task automation toward general-purpose intelligence. — From the positioning toward general-purpose physical AI

Summary: the real news is one model driving different bodies

What draws the eye in the Gemini Robotics 2 announcement is the footage of a humanoid walking around tidying up. But the heaviest technical change is a single fact: the same model checkpoint drove robots with different hand hardware, and a robot that is not a humanoid at all.

One reason robots have not spread is that each body needs its own rebuild, and the engineering invested does not carry over. If one model covers multiple bodies and adapts to a new one in hours with under 200 examples, that premise changes.

The limits DeepMind names are just as important. Multi-finger dexterity remains hard, and movement speed has room to improve. And the VLA model that actually moves the body is still early-access only. This is not technology entering your workplace now—it reads as groundwork that pays off over several years. What you can touch as of August 2026 stops at ER 2, the reasoning model.

How far Gemini has come as a language model is covered in our article on Gemini 3.6 Flash. Reading the two together makes the positioning clearer: this release is that progress flowing into a body.

Free ToolURL to Markdown ConverterConvert any public web page URL to Markdown. Preserves headings, tables, lists, and links — perfect for LLM and RAG preprocessing, research notes, and archiving web articles.Try it now →

FAQ

Q. What is Gemini Robotics 2?
A family of AI models for controlling robots, announced by Google DeepMind on July 30, 2026. The model at its center, Gemini Robotics 2, is a VLA (vision-language-action) model that converts camera input and spoken instructions into motor control, and DeepMind says it can drive a full humanoid from feet to fingertips.
Google DeepMind Blog — Gemini Robotics 2
Our most advanced vision-language-action model (VLA) that converts vision and language input into motor control, enabling a robot to take action. This model is capable of controlling full humanoids, from feet to fingertips, and other bi-arm robots. Google DeepMind Blog — Gemini Robotics 2
Q. How is this different from the previous Gemini Robotics?
The biggest difference is the range of the body it can control. Earlier models drove a humanoid's upper body for table-top work; Gemini Robotics 2 extends into whole-body motion such as walking, crouching and reaching.
Google DeepMind Blog — Humanoids in motion
While our previous models controlled the humanoid's upper-body to achieve table-top tasks, Gemini Robotics 2 expands physical AI into whole-body motions. Google DeepMind Blog — Humanoids in motion
Q. Can I try it right now?
Of the three models, the reasoning model Gemini Robotics ER 2 is available in Google AI Studio and in private preview on Gemini Enterprise Agent Platform. The VLA model that actually drives motors and the on-device model are limited to early-access partners, so they are not generally available.
Google DeepMind Blog — Gemini Robotics 2
Gemini Robotics ER 2, our reasoning model, is now available on Google AI Studio and in private preview on Gemini Enterprise Agent Platform. Our VLA and On-Device models are available to early-access partners. Google DeepMind Blog — Gemini Robotics 2
Q. What is it still bad at?
Plenty. In the caption of its own results chart, DeepMind states that while whole-body and gripper-based dexterous tasks reach a medium to high success rate, multi-finger dexterous manipulation remains challenging. It also acknowledges that movement speed has room to improve.
Google DeepMind Blog — results chart caption
While Gemini Robotics 2 achieves a medium to high success rate for whole-body and gripper-based dexterous tasks, the multi-finger dexterous manipulation remains challenging. Google DeepMind Blog — results chart caption

Related Tools

Related Tool Categories

Articles