Gemini Robotics reaches physical control
Google DeepMind announced Gemini Robotics — a vision-language-action model built on Gemini 2.0 — along with Gemini Robotics-ER for embodied reasoning, claiming a doubling on a generalisation benchmark against other such models.
Why it matters
This is the point at which a frontier general-purpose model family grew an action output, rather than a robotics group borrowing a vision-language model. In the atlas it also opens a line previously represented only by Gemini Robotics 2 of July 2026.
DeepMind states that on average Gemini Robotics more than doubles performance on a comprehensive generalisation benchmark compared with other state-of-the-art vision-language-action models, and that Gemini Robotics-ER achieves a 2x-3x success rate over Gemini 2.0 in end-to-end control. Training data came primarily from the ALOHA 2 bi-arm platform; the model was also shown on bi-arm Franka arms and on Apptronik’s Apollo humanoid. Confidence is medium here: the benchmark is described as a comprehensive generalisation benchmark without reference to an external, independently run test, and the set of models compared against is not itemised in the post.