Signal YCurated AI News
更新于 8/20 09:5261 个信源

Google DeepMind 发布 Gemini Robotics ER 2:用视频理解、任务编排与多机器人协作驱动机器人

通用 AI7/30 23:00Google DeepMind查看原文 ↗
摘要

Google DeepMind 发布 Gemini Robotics ER 2,一个为机器人设计的高级推理模型,通过视频理解、实时空间推理和多步任务规划,使机器人能够自我修正并协作完成复杂任务,即日起开放 API。

核心要点
  • Gemini Robotics ER 2 可理解视频流,实时跟踪任务进度,并在出现错误时自适应调整。
  • 支持多机器人协作,使多个机器人能在共享空间内协同完成单机器人无法完成的复杂工作流。
  • 在进度分类任务上达到 57.4% 准确率,优于前代 Gemini Robotics ER 1.6 及竞争前沿模型。
  • 时刻定位性能显著提升,能精确识别关键事件发生的视频帧(如停止倒咖啡的时机)。
  • 模型集成工具调用(如 Google Search)和用户自定义函数,可充当高级规划器并调用低层控制接口。
原文佐证
  • Gemini Robotics ER 2 represents a step change in powering robots with video understanding, task orchestration, and multi-robot collaboration — making it possible for robots to be more helpful in the physical world.
  • Gemini Robotics ER 2 achieves 57.4% accuracy on progress classification tasks, outperforming previous generation models and competing frontier models.
AI 洞察
该发布意味着机器人从简单程序执行向自主智能体迈出关键一步,实时自我修正能力将显著降低机器人部署中的故障成本。多机器人协作原语有望推动仓储、制造等场景的自动化升级,而开放的 API 生态将加速物理 AI 应用的创新与落地。