Google DeepMind released Gemini Robotics 2 on 30 July as a three-model control stack rather than a finished humanoid system. The vision-language-action model converts vision and language into motor commands that reach from feet to fingertips; the embodied-reasoning model plans multi-step sequences lasting several minutes and coordinates multiple robots; the on-device variant runs locally and adapts to new bodies. One checkpoint operates three distinct embodiments—the Apptronik Apollo 2 fitted with SharpaWave hands, the same Apollo 2 with Inspire hands, and a Franka Duo with a Robotiq gripper—without separate policies for each form [E1].
Published internal evaluations expose the uneven boundary of that shared policy. On the Apollo 2 with Inspire hands, floor pickup succeeds 45.7 percent of the time, table pickup 68.4 percent, and shelf pickup 76.3 percent. Multi-finger tasks on the SharpaWave hand swing wider still: unscrewing a light bulb reaches 92 percent while sweeping debris into a dustpan falls to 32 percent, with bag-tying and ziplock sealing sitting between 40 and 44 percent [E2]. Gripper work on the Franka platform lands higher, between 74 and 90 percent, yet the spread itself marks where contact-rich recovery remains unreliable [E2].
Availability follows the same hierarchy. Gemini Robotics ER 2 is offered in public preview through Google AI Studio and the Gemini API, with a private enterprise tier; the full vision-language-action model and the on-device variant stay limited to early-access partners and trusted testers [E3]. No independent laboratory has yet replicated the full suite of whole-body success rates under physical conditions outside DeepMind’s own controlled evaluations [E4].
Demonstrations still show coherent locomotion and manipulation. Apollo 2 walks, crouches, reaches and places a watering can from a natural-language prompt, while two robots collaborate on a garage tidy under the reasoning model’s direction. Adaptation claims state that new bi-arm embodiments can be brought online with fewer than 200 examples and a few hours of data [E1]. DeepMind itself notes that multi-finger dexterity remains challenging and that movement speed needs improvement [E1].
The safety layer adds an open ASIMOV-Agentic benchmark that tests refusal of unsafe tool calls, uncertainty detection and requests for human intervention. ER 2 records gains on human-proximity and constraint-following tests relative to earlier releases, yet the model cards explicitly advise against safety-critical uses such as healthcare or transport [E5]. Those gains sit on the reasoning model alone; the action models that actually move the joints remain behind the access gate [E3].
Taken together, the same-checkpoint transfer advances the frontier of policy generality across morphologies, yet the published success table supplies the strongest counter-case: a system that removes a bulb nine times out of ten still drops the dustpan seven times out of ten. The rates are company-reported, the physical replications absent, and the public surface limited to the reasoning component. Progress is therefore measurable and bounded by the numbers DeepMind chose to print [E2][E4].
The policy is intended to drive several robot bodies; the table still records where it drops the object.