See like a Robot: Robot-Centric Pointmaps for Vision-Language-Action Models

  • 类型:arxiv
  • 标识:2607.11498
  • 链接:https://arxiv.org/abs/2607.11498
  • 主分类:multimodal
  • 形态:method
  • 被引:0
  • 被引来源:Semantic Scholar
  • S2被引:0
  • 影响力被引:0
  • TLDR:Pointmaps provide robot-frame 3D geometry while preserving the dense H x W grid expected by pretrained 2D VLAs, so they integrate into existing VLAs with minimal architectural change and improve both pi0.5 and SmolVLA and outperform representative camera-viewpoint and 3D-aware baselines.
  • 待LLM分类:否
  • 标题中文:像机器人一样看:面向视觉-语言-动作模型的机器人中心点图
  • TLDR中文:Pointmaps 在保留预训练 2D VLA 所需 H × W 稠密网格的同时,提供机器人坐标系下的 3D 几何信息,能以极小的架构改动集成到现有 VLA 中,并提升 pi0.5 与 SmolVLA 的性能,优于代表性的相机视点和 3D 感知基线。
  • 来源文件
  • /inbox/tom/_candidates/2026-07-21-rag-retrieval-reranking-candidates.json
  • /inbox/tom/_candidates/2026-07-21-agent-rag-longcontext-candidates.json
  • [S2 enrich]