Refacto AI

Podcast episode

One Brain, Any Body: Google DeepMind's Keerthana on Gemini Robotics 2, Cross-Embodiment & Humanoids

agents engineering inference open-weights tool-use

Google DeepMind shipped Gemini Robotics 2, a suite of three models for controlling physical robots, and Keerthana Gopalakrishnan, the lab's research lead on the project, joined Nathan Labenz and Erik Torenberg on The Cognitive Revolution to walk through what's real and what isn't.

The most useful thing Gopalakrishnan says is that robotics is still in its "GPT-2 era," meaning the field is functional but nowhere near general. Two numbers back that up: Nathan Labenz clocked six seconds from prompt to robot command through the live API, which she agreed kills factory-line use cases, and the model's memory holds roughly three minutes of what the robot has seen. One of the three models, the reasoning planner, is on a public API today. The model that actually moves the robot body is gated to three named hardware partners: Boston Dynamics, Apptronik, and Agile Robots.

You can build a robot that thinks and calls tools this week. You cannot build one that moves well unless Google invites you in. That gap is the actual product.

Full analysis

Google DeepMind just shipped Gemini Robotics 2, a suite of three models for controlling physical robots, and put one of them (the reasoning brain) on a public API today. The lab's own research lead, Keerthana Gopalakrishnan, says robotics is still in its "GPT-2 era." That matters more than the demo reel. When the person building the thing tells you it's early, believe her, and read the gap between what the videos show and what she admits is unsolved.

How hard is this to undo? For most readers, there's nothing to undo. You're not buying a humanoid this quarter. But the API is live, so a robotics or physical-automation team could start wiring it in this week. What's actually being decided is whether the "one brain, any body" pitch is real yet, or whether it's a research preview with good PR. Nothing sets a deadline here. The API will still be there next quarter, probably cheaper.

The Skeptic. Watch the two numbers Gopalakrishnan didn't bury. Nathan Labenz measured six seconds from prompt to robot command through the API. Six seconds. She agreed that's useless for a factory line. And the reasoning model's memory holds about three minutes of what the robot has seen. Three minutes. A robot that forgets everything older than a pop song is not running your warehouse unsupervised. Then the anecdotes: a robot put a basket over someone's head, another gave an unwanted massage at a China conference. Those aren't bloopers, they're the edge cases that decide whether a machine is safe near a human. The videos tie trash bags. The failure stories tell you where it actually is.

The Researcher. The "GPT-2 era" line is a precise claim, not modesty. She means two things are missing. First, few-shot learning: teaching a robot a new task from one or a handful of examples, the way you teach a language model with a few lines in the prompt. It doesn't work reliably across enough tasks yet. Second, cross-embodiment: a model trained on one robot body still performs poorly on a different body. Language models run identically on any hardware. Robot brains don't. The 200-demonstration figure says it plainly: even with a foundation model that "knows" physics, you still hand-feed roughly 200 task demos per new robot body. That's the distance to "generic brain," stated in a number.

The Builder. What ships Tuesday? Only one of the three models is actually open to you: Gemini Robotics ER 2, the reasoning planner, on Google AI Studio with standard tool-calling, the same interface you already use for digital agents. Boston Dynamics is running it on Spot for instrument reading and handing out snacks. That's real, and it's a clue: the live use cases are slow, async, inspection-style jobs where six seconds doesn't hurt. The part that actually moves the robot body (the VLA, the model that turns "pick that up" into joint movements) is trusted-tester only, limited to Apptronik, Agile Robots, and Boston Dynamics. So you can build a robot that thinks and calls tools today. You cannot build one that moves well unless Google lets you in.

The Compute Pragmatist. The structure here rhymes with the LLM market, and that's the useful part for buyers. Gopalakrishnan expects a Flash-and-Pro split in robotics: big cloud models for the hard reasoning, small distilled models running on the robot's own chips for speed and for when the network drops. That's the exact cost ladder you already pay on language APIs. Big model to discover the capability, cheap small model to deploy it. The reasoning brain runs on Gemini 3.5 Flash, which is already the cheap tier. The on-device model exists precisely because a robot in a warehouse with spotty Wi-Fi can't round-trip to the cloud for every decision. Plan for hybrid: cloud for thinking, edge for reflexes.

The Open-Source Advocate. Here's what's quietly damning. This is the most closed release in a field that publishes a lot. The model that does the interesting thing, controlling the body, is gated to three named hardware partners. No weights, no broad access. Meanwhile the contrast with China is worth naming plainly: the "Robot Olympics" robots that out-sprinted human record holders got the viral attention, and Gopalakrishnan correctly says foot speed is close to irrelevant because flat-ground running is easy to fake in simulation. But that same Chinese hardware ecosystem is churning out cheap robot hands fast. The Woojin hand has roughly a 10-year-old's grip; the SCHUNK hand lifts about 20 kg and opens jars. The brains may be gated. The hands are becoming a commodity, and not from a US lab.

The biggest disagreement is between the Builder and the Skeptic, and it's about what "available now" buys you. The Builder can integrate the reasoning brain today. The Skeptic says the six-second latency and three-minute memory mean today's real jobs are narrow: inspection, reading dials, slow handoffs. Both are right, and that's the actual state of play. Embodied AI you can call from an API exists. Embodied AI that earns its keep on a production line does not, yet.

The second tension is Researcher versus the hype cycle. The demo shows whole-body humanoid control that didn't exist a year ago, which is genuinely fast. But "200 demos per new robot body" and "doesn't transfer across embodiments" mean every humanoid deployment carries a per-platform adaptation bill. Anyone budgeting for robots on the assumption of zero-shot, buy-it-and-it-works transfer is budgeting for a product that isn't shipping.

What this hinges on: does cross-embodiment transfer actually improve, so one brain runs many bodies without 200 demos each time? That's the single fact that turns robotics from GPT-2 into GPT-3. Until it moves, every humanoid is a custom integration, and the economics stay ugly. The council leans skeptical on timing and impressed by pace. For nearly every reader, the right move this quarter is to treat Gemini Robotics ER 2 as a toy to understand, not a system to deploy, unless your use case is genuinely slow and inspection-shaped. Prediction: Google DeepMind's Gemini Robotics body-control model (the component that directly commands robot limbs) will remain gated to trusted hardware partners with no public API or open weights when DeepMind ships its next major Gemini Robotics update, by April 2027.

Confidence: Medium — the unsolved safety and liability problems Gopalakrishnan named have not gotten shorter since launch.

Why: When DeepMind released Gemini Robotics, Gopalakrishnan put the body-control layer behind a three-partner trusted-tester wall while opening only the reasoning layer publicly, and she tied that choice directly to physical safety incidents (a basket dropped over a head, an unwanted contact) and the liability of a 200-pound machine operating near people. A model that moves limbs carries physical liability that a text API does not, so labs open the planner and hold the actuator. For this to change, DeepMind would need to solve the cross-embodiment and few-shot adaptation problems Gopalakrishnan called unsolved and shrink the liability exposure substantially, and nothing in the current release suggests either is close by April 2027.

Revisit by 2027-04-08: We're right if the GR2 body-control model is still trusted-tester or partner-only with no general API or open weights. We're wrong if DeepMind opens the body-control model to general developers via public API or releases its weights before that date.

Comments