← Back to Insights

Gemini Robotics 2 Adds Whole-Body Control, but Speed and Finger Dexterity Remain Limited

Nils Liu
AI Google DeepMind Robotics Gemini News

TL;DR

Google DeepMind introduced Gemini Robotics 2, ER 2, and On-Device 2 for whole-body control, longer task planning, and local execution, while acknowledging limits in speed and multi-finger dexterity.

Gemini Robotics 2 Adds Whole-Body Control, but Speed and Finger Dexterity Remain Limited

The gap between a polished demonstration and a deployable robot can be tested over the next six months. If partners reproduce stable whole-body task success on new hardware with less than 200 examples, while publishing cycle time and human-intervention counts, Google DeepMind’s rapid-adaptation claim will have commercial weight. If the evidence remains limited to edited demonstrations, outsiders will still be unable to assess reliability.

Google DeepMind released Gemini Robotics 2 on July 30, 2026. The vision-language-action model converts visual and language inputs into motor control, extending from a humanoid robot’s feet to its fingertips. The company showed an Apptronik Apollo 2 walking to a table, collecting a watering can, and placing it on a specified shelf. Other demonstrations include crouching, stretching, and tidying a room. The Verge independently reported the change from the previous model’s upper-body focus to whole-body control, but the observed evidence still originates in videos and descriptions supplied by Google.

Motion, planning, and local execution are separated

The release has three model layers. Gemini Robotics 2 handles physical action. Gemini Robotics ER 2 observes the environment, decomposes instructions, and tracks tasks lasting several minutes and involving hundreds of decisions. Gemini Robotics On-Device 2 runs locally when an internet connection is unavailable or network latency is unacceptable. ER 2 is available through Google AI Studio and in private preview on the Gemini Enterprise Agent Platform. The two action models are restricted to early-access partners, so most companies cannot yet obtain the complete system for independent testing.

Google also demonstrated cooperation between different robots, with Apollo 2 directing a dual-arm robot to place tools in a bin. Separating high-level planning from low-level control lets ER 2 revise a plan after a failed step. The announcement does not disclose communication-failure rates, total task duration, or the number of human takeovers. A single successful sequence therefore cannot be converted into warehouse, factory, or household throughput.

A 22-degree-of-freedom hand remains a performance boundary

Gemini Robotics 2 can control a five-finger SharpaWave hand with 22 degrees of freedom on Apollo 2, performing actions such as sealing a bag and tying a knot. It also controls two-finger grippers on a Franka Duo for tightly packing objects. Google’s chart describes medium-to-high average success across groups of tasks, but the announcement does not provide a complete numerical result for every task. The company explicitly says multi-finger dexterous manipulation remains challenging and movement speed needs improvement. Both constraints affect hourly output and safe working distance; task completion alone does not establish useful cycle time or the distribution of failures.

The adaptation claim for On-Device 2 is more specific. Google says a new dual-arm embodiment can typically be adapted with a few hours of data and less than 200 examples, even when shape, sensors, and degrees of freedom differ substantially. This remains a result under company-controlled conditions. The sources do not specify the cost of selecting examples, training compute, whether failures are included, or how well the same data covers changes in lighting, objects, and location.

For safety, ER 2 can detect nearby people, invoke safety tools, and stop a robot when someone approaches too closely. Google also introduced the ASIMOV-Agentic benchmark for unsafe-tool refusal, feasibility assessment, uncertainty handling, and escalation to a human. Google calls ER 2 its safest robotics model so far on safety-constraint and human-proximity tests. The Verge confirms that these functions were announced, but it offers no third-party certification or field-incident record that would independently validate the safety claim.

Over the next three to six months, the useful evidence will be unedited partner trials with per-task success rates, completion times, emergency stops, and human takeovers. Commercial availability and pricing for ER 2 beyond private preview will matter as well. The architecture now connects whole-body control, longer-horizon planning, and offline execution; its measurable limits remain multi-finger manipulation, movement speed, and an undisclosed failure rate in production environments.

Sources:

Get the latest insights

Join the newsletter to receive my latest articles on GenAI, AI Agents, and architecture.

No spam. Unsubscribe anytime.