Google DeepMind introduced Gemini Robotics 2, a suite of three AI models that gives humanoid robots coordinated whole-body control for the first time. One model is already free to test in Google AI Studio, opening the door for student developers.
Google DeepMind on July 30 unveiled Gemini Robotics 2, a new generation of AI models designed to serve as the intelligence layer inside physical robots — not just controlling arms and grippers, but managing an entire humanoid body from legs to fingertips. The release marks the first time the company’s robotics models can handle locomotion, whole-body coordination, and dexterous hand movements under a single learned policy. Earlier versions were limited to upper-body control.
The announcement introduces three distinct models rather than a single system. Gemini Robotics 2 is the flagship vision-language-action (VLA) model that translates visual and verbal input into physical motor commands across a full humanoid, including walking, crouching and reaching. Gemini Robotics ER 2 is an “embodied reasoning” model that acts as the robot’s high-level planning brain — processing user instructions, breaking tasks into steps, tracking progress, and coordinating with the VLA to carry out each action. Gemini Robotics On-Device 2 is a leaner VLA built to run locally on the robot itself, without requiring a network connection, and can adapt to an entirely new robot body in a few hours using fewer than 200 training examples.
Hardware partners for the launch include Apptronik, whose Apollo 2 humanoid was demonstrated picking up a watering can, walking to a storage shelf, and placing it precisely in a designated bin. The same model was tested with two different hand types — SharpaWave’s five-fingered, 22-degree-of-freedom hand and Inspire hands — as well as on the Franka Duo bi-arm platform with a Robotiq gripper. The On-Device model was also demonstrated on smaller and lower-cost bi-arm systems including the Dexmate, Trossen and SO101 research robots.
A new multi-robot collaboration capability lets different robot types communicate and divide up work on tasks too complex for a single machine. Alongside the models, DeepMind introduced ASIMOV-Agentic, a safety benchmark designed to measure whether a robot agent can refuse unsafe commands, predict whether a task is actually feasible, and ask a human for help when it isn’t sure.
Where This Fits in a Crowded Race
Google’s strategy stands apart from most rivals. Rather than building its own humanoid hardware, DeepMind is positioning its models as the intelligence layer that other manufacturers can plug into their robots — a platform play rather than a vertically integrated product. That’s a direct contrast to Figure AI, which cut ties with OpenAI in early 2025 and built its own in-house VLA model, Helix, to run on its Figure 03 hardware. Tesla takes a similar own-the-whole-stack approach with its AI-derived control networks and custom chips. Unitree, meanwhile, open-sourced a research VLA model in early 2026.
Nvidia ccupies a different lane, anchoring physical AI infrastructure through its Cosmos world foundation model, Isaac GR00T robot learning model, and Isaac Sim simulation environment, with partnerships across more than 150 robotics companies. Startup Physical Intelligence, which raised a $600 million Series B in November 2025 at a $5.6 billion valuation, is building a single general-purpose model meant to control any robot for any task. None of these players has yet established a clear lead. The physical AI foundation model market is still in early formation, and Google’s platform bet is as plausible as any other approach.
What Students and Developers Can Do Right Now
The most immediately usable piece of the release is Gemini Robotics ER 2, the reasoning and planning model, which is available now through the Gemini API and Google AI Studio. Google AI Studio is free — no subscription, no seat fees — meaning students can start building and testing robot planning agents today at no cost. The model accepts multimodal inputs including streamed video, audio and text, and developers can declare low-level control interfaces, such as navigation or manipulation APIs, as tools the model can call.
For students in robotics, mechatronics or computer science who already own or have access to low-cost robot arms — including the SO101 and Trossen platforms that DeepMind itself demoed — this is a practical on-ramp to working with frontier AI. Google has published sample notebooks on GitHub for wiring the model to robot control interfaces. The VLA and On-Device models are currently restricted to early-access partners, so full humanoid experimentation isn’t open to the public yet, but the reasoning layer is genuinely accessible.
Career-wise, the announcement reinforces a trend worth tracking. Google, Nvidia, Physical Intelligence, Figure AI and a wave of well-funded startups are all actively staffing embodied AI teams. The ability to bridge large foundation models with physical hardware — understanding both the ML side and the robotics control side — is shaping up as a distinct and high-demand engineering skill. For students still choosing a specialization, that gap between AI and physical systems is one of the more interesting places to plant a flag right now.
Source: Google DeepMind
Additional research sources
- https://www.bloomberg.com/news/articles/2026-07-30/google-unveils-gemini-ai-for-robots-struggling-with-dexterity
- https://blog.robozaps.com/b/gemini-robotics-2-humanoid-robot-ai
- https://ai.google.dev/gemini-api/docs/robotics-overview
- https://blog.google/innovation-and-ai/models-and-research/google-deepmind/gemini-robotics-er-2/
- https://the-gadgeteer.com/2026/08/01/google-gemini-robotics-er-2/
- https://blog.robozaps.com/b/best-humanoid-robots
