EyeRobot 2.0: Active Gaze for Precise Manipulation without Wrist Cameras

  • 类型:arxiv
  • 标识:2610.03710
  • 链接:https://arxiv.org/abs/2610.03710
  • 主分类:multimodal
  • 形态:method
  • TLDR:Inspired by human vision, we introduce a framework using active gaze to enable fine-grained bimanual manipulation with only a single stereo camera. EyeRobot 2.0 physically attends to a 3D fixation point in the scene by swiveling two eye viewpoints to center their gaze on it. The resulting images are processed foveally by allocating more visual tokens to the image centers, focusing computation on task-relevant features. Such Active Visual Fixation (AVF) requires carefully coordinated gaze during task execution, which we accomplish hierarchically by first training a low-level gaze servoing polic
  • 待LLM分类:否
  • 标题中文:EyeRobot 2.0:无需腕部摄像头的精确操作主动注视
  • TLDR中文:受人类视觉启发,本文提出一种框架,通过主动注视仅用单个立体相机即可实现细粒度双臂操作。EyeRobot 2.0 通过旋转两个视角使其注视中心对准场景中的 3D 注视点,从而物理性地关注该 3D 注视点。所得图像以中央凹方式处理,为图像中心分配更多视觉 token,将计算聚焦于任务相关特征。这种 Active Visual Fixation (AVF) 需要在任务执行过程中精心协调注视,我们通过分层方式实现:首先训练一个低级注视伺服策略
  • 来源文件:
  • /inbox/tom/_candidates/2026-10-06-rag-retrieval-reranking-candidates.json
  • /inbox/tom/_candidates/2026-10-06-agent-rag-longcontext-candidates.json