02—Gemini Robotics ER 2 Is the High-Level Brain, Not the Motor Controller
Google describes Gemini Robotics ER 2 as a high-level brain for robots.
That wording defines its place in the stack.
The model receives text, images, video and audio.
It reasons about the physical environment.
It can choose tools.
It can plan a sequence.
But the lower-level movement can remain inside another system.
A VLA model can translate vision and language into motor action.
A navigation API can move a mobile robot.
A manipulator API can control an arm.
A grasping system can handle the final contact with an object.
ER 2 coordinates those capabilities instead of replacing every controller underneath them.
This creates a hierarchy.
The user states the goal.
ER 2 decides what the goal requires.
The robot APIs or VLA models execute the physical operations.
That separation also lets the same reasoning architecture work with different robot bodies.
The model can call whatever lower-level interface the developer exposes as a tool.
03—Robot Capabilities Become Tools the Model Can Orchestrate
Gemini Robotics ER 2 uses a tool-oriented architecture.
Developers can expose low-level robot controls as callable functions.
A navigation function can move to a location.
A camera function can inspect the scene.
A manipulator function can position an arm.
A VLA model can perform a learned physical behavior.
The reasoning model does not need every physical detail encoded inside one prompt.
It can select from the available tools as the task progresses.
Google’s developer documentation calls this task orchestration.
The model combines spatial reasoning with custom robotics APIs to complete longer sequences.
That makes robot control look increasingly like software-agent orchestration.
The difference is the output.
A software tool may edit a file.
A robotics tool may move a machine.
The planning structure is similar: understand the current state, choose the next capability, receive the result and continue.
04—Continuous Video Gives the Planner a Live View of Progress
ER 2 adds a stronger temporal layer through continuous video.
A physical task unfolds over time.
The robot moves.
An object changes position.
A container becomes full.
A tool reaches the intended location.
The reasoning system needs to know when one step is actually complete before it moves to the next.
Google says ER 2 can watch continuous video feeds to track progress and verify task completion.
That gives the planner a live source of state rather than only a sequence of isolated images.
The model can observe what the robot is currently doing while it is already reasoning about the following step.
This changes the rhythm of the control loop.
Perception and planning can overlap.
The robot does not always have to stop completely before the reasoning layer begins the next decision.
The model can keep watching the task as it unfolds.
05—Moment Finding Helps the Model Decide When a Step Is Complete
One of the new evaluation areas is moment finding.
The model receives a video and has to identify the moment when a defined event occurs.
In robotics, that can mean the moment a subtask has been completed.
Google gives examples such as tightening a light bulb or tying a trash bag.
ER 2 can use the continuous video to decide when the required state has been reached.
Google reports 91.3 percent accuracy on its moment-finding evaluation with a mean absolute distance of 0.96 seconds.
Google also reports that the model executes this evaluation four times faster than the larger comparison category used in its release.
Those are Google’s published benchmark results for its test setup.
The important architectural point is that timing becomes part of embodied reasoning.
The model is not only asking what is in the scene.
It is asking when the scene has reached the state required for the next action.
06—Planning Can Continue While the Robot Is Already Acting
Google designed the ER 2 control loop so reasoning and action can overlap.
The model can think about the next step while the current physical action is still being performed.
This matters because robotics operates in real time.
A long pause between every decision can change the feel of a multi-step workflow.
ER 2’s streaming version integrates with the Gemini Live API through a bidirectional streaming endpoint designed for latency-sensitive robotics work.
The streaming endpoint can receive continuous audio and video while using function calling to control robot tools.
That gives the planning system a persistent connection to the physical process.
The robot acts.
Video continues arriving.
The model updates its understanding.
The next tool call can be prepared from the current state.
This is closer to an ongoing control conversation than a sequence of isolated prompt-response requests.
07—Google Provides Two ER 2 Endpoints for Different Control Loops
Gemini Robotics ER 2 currently has two model endpoints.
The standard preview endpoint is gemini-robotics-er-2-preview.
It supports multimodal input, function calling, structured outputs, code execution, search grounding and other Gemini capabilities.
The streaming endpoint is gemini-robotics-er-2-streaming-preview.
It is optimized for the Live API and low-latency robotics workflows.
Both accept text, images, video and audio as input.
Both produce text output that can include reasoning results and tool calls interpreted by the application.
Google documents an input context window of 131,072 tokens and an output limit of 65,536 tokens for the current preview endpoints.
The two-endpoint structure gives developers a choice.
A task that does not require continuous low-latency control can use the standard model.
A robot that needs ongoing video and audio interaction can use the streaming path.
08—Multi-Robot Collaboration Adds Another Level Above One Machine
The largest conceptual change in ER 2 is multi-robot collaboration.
Google demonstrates different robot bodies working on one shared task.
The reasoning layer can understand that the machines have different physical capabilities.
It can then coordinate a handoff between them.
Google’s release shows Apptronik’s Apollo 2 working with a Franka FR3 Duo system.
The two robots do not need identical bodies.
One can be better positioned for one part of the workflow.
Another can handle the next part.
ER 2 provides the shared semantic plan between them.
This turns robot diversity into something the planner can reason about.
The fleet does not need to behave as several copies of one machine.
Different forms can contribute different capabilities to the same objective.
10—Apollo 2 Shows How a Humanoid Can Become One Capability in the Team
Apptronik’s Apollo 2 is one of the robot platforms shown in Google’s multi-robot work.
Apollo 2 is a modular humanoid platform available with bipedal and wheeled-base configurations.
Apptronik describes it as a physical platform designed for embodied AI, with perception, manipulation and mobility systems that can work with higher-level models.
Apptronik and Google DeepMind already have a research partnership around Gemini Robotics.
Apptronik’s Robot Park also collects real-world task data from Apollo 2 fleets for the development of future robotics models.
Inside a multi-robot plan, the humanoid does not need to become the entire system.
It can be one physical participant.
The reasoning model can assign the part of the task that fits Apollo’s body and then coordinate another machine for a different step.
11—Franka's Dual-Arm Platform Adds a Different Physical Skill Set
The Franka FR3 Duo represents another kind of robot.
It uses two collaborative robotic arms rather than a humanoid body.
Franka has demonstrated Gemini Robotics with its dual-arm systems for real-time planning and multi-step manipulation.
This gives the shared planner another physical form to work with.
A dual-arm station can specialize in coordinated manipulation around a fixed workspace.
A mobile or humanoid platform can approach the task from another position.
The high-level model can reason about the complete workflow while the local robot controller handles the detailed motion of each arm.
That is the advantage of separating embodied reasoning from body-specific execution.
The reasoning model can remain common.
The machines underneath can stay specialized.
12—Spot Shows the Same Planning Pattern Through Robot APIs
Boston Dynamics Spot demonstrates the same architecture from another direction.
Google used ER 2 to orchestrate Spot APIs for navigation and manipulator movement.
The user can give a natural-language task.
ER 2 can decide which Spot capabilities need to be called and in what order.

Boston Dynamics has also documented earlier work connecting Gemini Robotics to Spot through a tool layer built on the Spot SDK.
The reasoning model acted like a high-level operator.
Spot’s existing autonomy, navigation and manipulation systems remained underneath it.
ER 2 extends that orchestration model with real-time streaming and stronger progress understanding.
The important part is that the robot API becomes a tool interface.
A mature robot platform does not have to discard its own control stack.
The higher-level model can coordinate the capabilities that already exist.
13—Spatial Reasoning Connects Language to Physical Coordinates
Planning in the physical world also requires spatial grounding.
Google’s robotics API documentation exposes capabilities for pointing to objects, tracking them in video, detecting bounding boxes and planning trajectories.
A model can receive an image and return normalized coordinates for visible objects.
Those coordinates can then be passed to another robotics system.
This creates a bridge between language and geometry.
The user says which object matters.
The model identifies it.
A robot controller receives the spatial output.
The physical system acts on that location.
Spatial reasoning becomes another shared service in the planning layer.
It is not tied to one manipulator.
Any compatible lower-level system can use the structured location information the model produces.
14—Tool Calls Can Reach Software Services as Well as Robot Hardware
ER 2 can also call non-robot tools.
Google says the model can use services such as Google Search or user-defined functions.
That widens the planning loop.
A robot task may require information that is not visible in the room.
The model can retrieve the information.
Then it can continue the physical workflow.
The same agent can therefore move between digital and physical tools.
Search can answer a question.
A database can provide a parameter.
A robot API can move a machine.
A VLA model can execute manipulation.
ER 2 coordinates the sequence.
This is another reason embodied reasoning is becoming part of the broader agent architecture.
The distinction between software tools and physical tools remains important at the execution layer.
At the planning layer, both can appear as capabilities available to the same task.
15—ER 2 Is Available to Developers as a Preview
Gemini Robotics ER 2 is currently distributed as a preview model.
Google makes the standard endpoint available through the Gemini API and Google AI Studio.
Google also lists Gemini Enterprise Agent Platform as a private-preview distribution path.
The model card identifies ER 2 as a Vision-Language Model based on Gemini 3.5 Flash.
Google says it was trained on Gemini 3.5 data together with additional embodied-reasoning datasets.
The current release is therefore a developer platform as well as a research result.
Developers can connect their own robot APIs.
They can stream multimodal input.
They can test spatial reasoning, progress understanding and task orchestration.
The preview status is important because it defines where the technology sits today.
The architecture is accessible.
The ecosystem is still actively developing around it.
16—Robots Are Starting to Plan Together, Not Just Act Alone
Gemini Robotics ER 2 shows a different direction for robotics.
The intelligence layer does not have to live entirely inside one robot body.
A high-level model can watch the environment.
Understand the shared task.
Plan several steps.
Track progress through continuous video.
Decide when one step is complete.
Call robot APIs or VLA models.
Coordinate different machines.
Keep reasoning while physical actions continue.
Apollo 2, Franka’s dual-arm platform and Spot illustrate three different robot forms that can participate in this kind of architecture.
The bodies stay different.
Their lower-level controllers stay different.
The planning layer provides a common semantic structure above them.
That is the important shift.
The future robot team may not be a fleet of identical machines running the same behavior.
It can be a collection of specialized bodies coordinated by a shared reasoning system.
The robot stops being the only unit of intelligence.
The workflow becomes the unit.
That is the upgrade.
