THE APEX TIMES
Alphabet’s Google DeepMind launches Gemini Robotics ER 2, a “high-level brain” for real-world robot video tasks
The new Gemini Robotics ER 2 model is designed to orchestrate multi-step actions from continuous video, call tools in real time, and coordinate work across multiple robots.
Alphabet’s Google DeepMind on Wednesday unveiled Gemini Robotics ER 2, a robotics-focused AI model aimed at improving how robots interpret video, plan multi-step tasks, and execute those plans without getting stuck. The company positions the system as a “high-level brain” for robots that can reason in real time from what it sees, decide what to do next, and then hand off the physical act of moving and manipulating objects to lower-level control systems.
Gemini Robotics ER 2 is built for embodied reasoning, meaning it is meant to operate at the speed of the physical world rather than waiting for slow, step-by-step deliberation. Google says the model integrates tool orchestration and multi-robot collaboration, and that it can keep working without “jarring stop-and-think pauses” by streaming commands through Google’s Gemini Live API, which uses a bidirectional endpoint designed for low latency tasks.
A core feature of the launch is the model’s ability to follow task progress from continuous video. Google describes progress classification as a framework that assigns each video frame to one of five completion bands (0 to 20%, 20 to 40%, 40 to 60%, 60 to 80%, and 80 to 100%). With that quantification, the robot can adjust on the fly, retry failed steps, and decide when it is time to move to the next stage of a workflow. In Google’s evaluations, it reports 57.4% accuracy on progress classification tasks.
Google also highlights “moment finding,” which it describes as identifying the exact video frame where a key event occurs, such as determining the correct moment to stop pouring coffee into a cup. For moment-finding tasks, Google says Gemini Robotics ER 2 achieves 91.3% accuracy and a 0.96-second mean absolute distance, emphasizing that safe physical robotics needs sub-second latency.
The company says Gemini Robotics ER 2 improves tool orchestration compared with Gemini Robotics ER 1.6 across three different control modes: real-world VLA (vision-language-action) models, simulated VLA, and human tele-operation (remote control by a person). Google does not provide the absolute performance deltas in the material, but it says the newer model “consistently outperforms” the prior version for orchestrating tools used by robotic systems.
To illustrate how the system can be used, Google points to a demo involving Boston Dynamics’ Spot robot. In that example, Gemini Robotics ER 2 is used to orchestrate Spot APIs for capabilities like navigation and manipulator movement so the robot can complete a user command in an interactive setting. Google says the demo is framed around Spot fetching objects, such as a popcorn snack, when prompted in natural language.
Beyond single-robot use, Google says Gemini Robotics ER 2 enables multi-robot collaboration in shared spaces. The company describes this as letting different robots communicate using a shared semantic understanding, allowing handoffs and coordination on tasks that would be difficult for one machine alone. Google names two partner systems in the launch material, Apptronik’s Apollo 2 and Franka’s F3 Duo, as examples it claims can collaborate through the model.
Google also presents safety as a central part of the release. In its description, Gemini Robotics ER 2 is positioned as its safest robotics model, with gains on benchmarks that evaluate safety instruction following and human proximity. The company says it observed the model halting a humanoid robot when a person is nearby and resuming only after the area is clear. It also introduces a benchmark intended to test whether a foundation model can act as a safe VLA orchestrator by enforcing safety constraints, monitoring the environment, checking physical feasibility, and seeking human clarification when needed, with more detail referenced in a safety technical report.
For developers and robotics builders, the company says Gemini Robotics ER 2 is now available via the Gemini API, Google AI Studio, and as a private preview on the Gemini Enterprise Agent Platform. Google describes an “agentic setup” where developers declare low-level control interfaces, such as VLA models or navigation APIs, as tools, then stream multimodal inputs (including video, audio, or text) directly into the model. The company also says it is sharing configuration and prompting examples to help users start building physical AI agents. What remains unclear from the public launch material is how developers should best balance compute cost, latency targets, and safety behaviors across different robot platforms and environments, since the release does not provide standardized deployment guidance across vendors.
The bigger business question for Alphabet is whether these capabilities translate into practical adoption by robotics companies and integrators. The launch directly targets developers building “helpful robots” for everyday settings, and it positions tool orchestration and real-time progress verification as the missing pieces that make robots more reliable rather than merely more capable. With Gemini Robotics ER 2 now accessible for development, the next phase will likely be measured by how quickly teams can integrate it into real deployments, and whether the model’s reported accuracy and safety behaviors hold up across a wider variety of tasks, sensors, and operating conditions.
Why It Matters
- Robotics systems increasingly depend on software that can reliably interpret video and decide what to do next; Google is targeting those failure points with progress tracking and moment detection.
- Low-latency orchestration is a commercial barrier for physical robots, and Google’s streaming design suggests an effort to make AI behavior more usable in live environments.
- If the model’s safety behaviors generalize, it could help reduce one of robotics’ core deployment blockers: unpredictable reactions around people.
- Multi-robot collaboration could expand market opportunities beyond single-machine demos, but the economic value will depend on integrators’ ability to coordinate heterogeneous hardware.
Key Facts
- Google DeepMind launched Gemini Robotics ER 2, a robotics-focused model it describes as a high-level “brain” for real-world robot tasks.
- The model is designed for real-time spatial reasoning and multi-step task planning, with tool calls and streaming orchestration via the Gemini Live API to reduce stop-and-think delays.
- Google says Gemini Robotics ER 2 improves task progress understanding using progress classification across five completion bands and moment finding to identify key event frames.
- In Google’s reported evaluations, progress classification achieved 57.4% accuracy, while moment-finding achieved 91.3% accuracy and a 0.96-second mean absolute distance.
- Google says Gemini Robotics ER 2 outperforms Gemini Robotics ER 1.6 for tool orchestration across real VLA, simulated VLA, and human tele-op control modes.
- The model is publicly available to developers via the Gemini API and Google AI Studio, with private preview access on the Gemini Enterprise Agent Platform.
Technology Related
AMD says Instinct AI systems are now operating in Saudi Arabia, highlighting a potential ramp tied to additional data-center power
A recent market report frames AMD’s Instinct deployments in Saudi Arabia as a move from plan to production, and points to how incremental data-center capacity, measured in megawatts, could influence investor expectations.
Salesforce says AI-driven revenue momentum is building as Agentforce adoption spreads
In a recent market update circulated by Yahoo Finance, Salesforce management pointed to expanding use of its AI offerings, including agentic workflows and consumption-style pricing, as the company positions its next growth phase.
Salesforce backs HiBob to bolster workforce AI, and adds a new AgentExchange email tool
Salesforce said it is supporting HR-analytics and talent-workforce platform HiBob as part of efforts to connect enterprise data with “powered AI.” The company also announced an AgentExchange email tool aimed at expanding what business agents can do inside everyday workflows.
EverPass Media expands NFL distribution via multi-year Netflix deal for 2026 slate
EverPass Media says it has added Netflix’s five NFL games for the 2026 season to its NFL distribution offering, including the first-ever Thanksgiving Eve game, plus “NFL Honors.”
Broadcom leans harder into VMware AI with a push aimed at enterprise rivals
Broadcom’s VMware AI push is tied to the latest VCF 9.1 release, as the company’s messaging positions it against Nutanix and Microsoft in hybrid cloud and enterprise AI rollouts.
Yahoo Finance points to “buy zones” for Microsoft, Palantir, Shopify and ServiceNow
A market-readout from Yahoo Finance flagged several software and AI-linked names, including Palantir (PLTR), as trading in or near so-called buy zones. The note is framed as technical or timing-oriented, with limited company-specific detail.
Oracle Shares Fall as Investors Focus on Cash Flow Gap and Rising Borrowing Costs
A reported $23.7 billion cash shortfall over Oracle’s last fiscal year and $43 billion in borrowing are drawing attention to the company’s interest-rate exposure, a factor that can quickly change sentiment when Treasury yields are elevated.
Adobe’s next report faces a split view: Citi still expects a beat, but flags lingering risks
After Adobe lowered its annual revenue outlook, one analyst said the company can still deliver a beat-and-raise in fiscal third-quarter results, even as concerns remain.
Palantir’s commercial growth may overtake government revenue sooner than expected, according to a new market model
A widely watched growth-math forecast argues Palantir’s commercial revenue could surpass its government revenue before 2027, driven by a widening gap in the companies’ growth rates.
Netflix shares face another round of debate after new market commentary, but company keeps details scarce
A recent Yahoo Finance-linked article argues Netflix is not finished telling its story, urging investors to stay cautious until more clarity emerges.