Google 发布 Gemini Robotics 2
VLA 模型 + 具身推理 + 端侧模型,人形机器人演示与安全基准同步
Google DeepMind 发布 Gemini Robotics 2:VLA 模型 + ER 2 具身推理 + On-Device 2 端侧模型,在 Apptronik Apollo 2 人形机器人上演示,并推出 ASIMOV-Agentic 安全基准。
2026 年 8 月 4 日,Google DeepMind 发布 Gemini Robotics 2。演示里,Apptronik 的 Apollo 2 人形机器人在真实环境里完成操作——不是预先编排的固定动作,而是看着场景、理解指令、规划步骤、再伸手执行。机器人第一次把「看懂」和「动手」放进同一条链路。
这套系统由三部分组成。VLA 模型负责视觉—语言—动作的映射,把「把红色杯子放到托盘上」这样的指令变成具体的关节运动;ER 2 具身推理模型处理长程任务规划,在多个步骤之间保持目标;On-Device 2 面向端侧低延迟部署,让机器人不必每次动作都等云端往返。三件套的分工,对应的是真实机器人系统的三个瓶颈:理解、规划、执行。
与模型同时发布的 ASIMOV-Agentic 安全基准,是这次发布里容易被忽略却分量最重的一块。过去具身智能的安全评估沿用通用 AI 基准,测的是「模型会不会说错话」;ASIMOV-Agentic 测的是「机器人在自主操作中会不会做错事」——碰撞、误操作、在不确定时是否知道停下。当机器人开始拿工具、进厨房、靠近人,安全不再只是论文里的章节,而是产品能不能上线的门槛。
演示视频需要带着边界读:Apollo 2 的展示是受控场景,不是量产承诺;VLA 模型也不是通用的机器人操作系统,换一个机械臂、换一套传感器,泛化仍要重新验证。但方向已经清楚:具身智能的竞争,从「谁能控制单任务」升级为「谁能提供理解、规划、执行、安全一整套栈」。
Gemini Robotics 2 没有让机器人一夜之间走进工厂和家庭。它做的是把「看懂场景再行动」变成可组合的模型家族,并第一次给具身 Agent 配了专门的考卷。当机器人开始为画面负责、为动作负责、为安全负责,人形机器人就从演示片走进了工程与评测的领域——这一步,比任何一次惊艳的 demo 都更接近量产。
On August 4, 2026, Google DeepMind released Gemini Robotics 2. In the demo, Apptronik's Apollo 2 humanoid performed operations in a real environment—not pre-scripted fixed motions, but looking at the scene, understanding the instruction, planning the steps, then reaching out to execute. For the first time, a robot put "understanding" and "acting" into one chain.
The system has three parts. The VLA model maps vision-language-action, turning "put the red cup on the tray" into concrete joint motions; ER 2 embodied reasoning handles long-horizon planning, keeping the goal across multiple steps; On-Device 2 targets low-latency edge deployment so the robot does not wait for a cloud round-trip on every move. The three-way split corresponds to the three real bottlenecks of robot systems: understanding, planning, execution.
The ASIMOV-Agentic safety benchmark released alongside the models is the easiest piece to overlook and the heaviest. Embodied-AI safety used to borrow generic AI benchmarks that test whether a model says the wrong thing; ASIMOV-Agentic tests whether a robot does the wrong thing during autonomous operation—collisions, misoperations, knowing when to stop under uncertainty. Once robots pick up tools, enter kitchens, and work near people, safety stops being a paper chapter and becomes the gate for shipping a product.
Demo videos need their boundaries read: the Apollo 2 showcase is a controlled scene, not a mass-production promise; the VLA model is not a general robot operating system—swap the arm or the sensors and generalization must be re-verified. But the direction is clear: embodied-AI competition upgraded from "who can control a single task" to "who can deliver the whole stack of understanding, planning, execution, and safety."
Gemini Robotics 2 did not put robots into factories and homes overnight. What it did was turn "understand the scene before acting" into a composable model family and give embodied agents their first dedicated exam. When robots become accountable for the picture, for their actions, and for safety, humanoids move from demo reels into engineering and evaluation—a step closer to production than any impressive demo.
展开完整事件档案人物、主题、模型与产品
- 人物
- —
- 模型
- gemini-robotics-2
- 产品
- —