Google DeepMind introduced Gemini Robotics 2.0, promising improved dexterity and safety over the previous generation. The most-discussed demonstration shows the model controlling a humanoid robot's entire body rather than a single manipulator arm.
The difference between controlling an arm and controlling a whole body isn't one of degree. A fixed arm has a stable base and a known working envelope. A body that walks has to solve balance, weight transfer and stumble recovery while performing the task, and each of those affects the others in real time.
The emphasis on safety in the launch material isn't incidental. A humanoid robot moving through space shared with people has a risk surface a caged arm doesn't, and it is terrain where a model error becomes a physical injury.
What remains open is what happens when the model meets a situation outside its training. In software, a model that doesn't know produces bad text. In hardware that weighs tens of kilos and walks near people, the same uncertainty has a different consequence.
The announcement consolidates the year's trend: language models have stopped being text generators and become a control layer for physical systems.
