
Robots don’t fail only because a hand misses a cup. They fail because the task around that cup keeps changing. Google DeepMind’s answer is to split the problem, and Gemini Robotics ER 2 is the half that keeps track of the larger job.
ER 2 isn’t a robot you can buy, and it isn’t the whole announcement either. It’s one of three models in the Gemini Robotics 2 release, and it sits above the hardware: it takes an instruction in plain language, breaks it into steps, and can reach out to Google Search or a developer’s own functions when it needs something it doesn’t already know.
What Google Announced
Google calls ER 2 “a high-level brain for robots” and shipped it alongside two siblings: Gemini Robotics 2, a vision-language-action model that controls full humanoids from feet to fingertips, and On-Device 2, a lightweight version that runs locally on the robot itself. Only one of the three is doing the planning.
ER 2 decides what should happen. A separate action model moves the limbs. Splitting the two is what lets Google aim the same reasoning at very different machines, from a stationary arm to a walking humanoid.
The model is built on Gemini 3.5 Flash and connects through the Gemini Live API, which Google says removes the stop-and-think pauses that make robot demos look stilted. The company frames the whole release as an upgrade over ER 1.6, the previous version, and most of its published comparisons use that model as the baseline.
Why the Continuous View Changes the Pitch
A robot that only checks isolated snapshots can lose the thread of a longer task. Google’s claim is that watching the job happen, instead of checking a still frame afterward, is what lets a robot catch its own mistake early enough to matter. For a kitchen, a warehouse, or a shared workspace, that matters more than another polished demonstration.
The numbers Google published are its own, not independent testing. ER 2 scores 57.4 percent on progress classification, which sorts each video frame into one of five completion bands, so it leads Google’s comparison set while still landing outside the right band more than four times in ten. On moment-finding, the job of spotting the exact frame where something happens, such as when to stop pouring, Google reports 91.3 percent accuracy and a 0.96-second mean absolute distance against the much larger models Google says it competes with, at a fraction of their compute cost and four times their execution speed. That distance is an error measure rather than a speed, so on average the frame it picks sits about a second from the real moment, and Google measured that on video rather than on a robot mid-pour.
A Model That Can Coordinate More Than One Machine
The headline addition is that ER 2 can direct more than one machine at once. Google’s argument is that some jobs need a rover and a humanoid rather than a better single robot, and that different machines can now share enough understanding to hand work back and forth. Its collaboration demo pairs Apptronik’s Apollo 2 humanoid with Franka’s FR3 Duo bi-arm robot, and a separate demo runs on Boston Dynamics’ Spot.
The coordination problem already shows up in commercial settings. When we looked at KEENON’s robot café, the takeaway was that a humanoid doing the visible work still needs a supporting cast of machines around it, and something has to direct the traffic between them.
That doesn’t mean a household robot crew is close. Google hasn’t attached ER 2 to a consumer robot or named a retail partner, and every machine in its demos is a research or industrial platform. The immediate audience is developers and businesses that already own hardware and have a reason to connect it.
The Safety Claim Is the Part Worth Watching
Google calls ER 2 its safest robotics model, and the specific behavior is what stands out: in its own testing, the model halted a humanoid when a person came near, then resumed on its own once the space was clear. Google hasn’t published a success rate for that. The company reports gains on two safety benchmarks, one measuring whether a model follows physical safety constraints and one measuring how well it detects people nearby.
Google also published a new benchmark for whether a model can safely supervise an action model, testing its capacity to enforce constraints, monitor its surroundings, judge what is physically feasible, and ask a human when it isn’t sure. That last behavior matters more than any accuracy figure. A robot that knows when to stop and ask is a different product from one that always tries.

Who Can Use It Now
If you write code, you can reach the model today through the Gemini API and Google AI Studio, though both endpoints are labeled preview rather than production. Enterprise access is still gated behind a private preview on the Gemini Enterprise Agent Platform. Anyone already building on ER 1.6 has a deadline, because Google says that model shuts down at the end of August and the upgrade amounts to swapping the model string.
Google published sample notebooks on GitHub for configuring the model and wiring it to robot control interfaces. Developers declare things like action models or navigation APIs as tools, then stream video, audio, or text straight into ER 2.
If you don’t write code, there’s nothing here to sign up for. No consumer pricing, no robot bundle, and no timeline for anything you would put in a house.
What Is Still Unclear
Google didn’t say which physical robots get ER 2 beyond its demo partners, what a real deployment costs, or how often a human has to step in when nobody is filming. Those three answers decide whether this becomes an operating layer or another impressive robotics video.
The projects that clear the gap between a planning layer and a working machine tend to be narrow and purpose-built, like the robot dog mobility chair a son built so his dad could get off-road again. Until Google shows intervention rates outside its own evaluations, ER 2 stays a capability announcement rather than a shipping product.
The TG Take
The interesting part isn’t that Google wants robots to move with more dexterity. Plenty of robotics demos already do that.
ER 2 is an attempt to give a robot a sense of the larger assignment while the room, the objects, and the other robots keep changing. That’s the part that has held this category back, and it’s harder than the hands.
If the reasoning layer holds up outside Google’s own evaluations, robot hardware may start looking less like the bottleneck. The safety behavior is the tell worth watching, because the day a robot reliably stops and asks is the day one belongs in a house.
Sign up for our newsletter today.
No ads, no spam, just links to our latest articles!
Auto Amazon Links: No products found.















