See like a Robot: Robot-Centric Pointmaps for Vision-Language-Action Models
- 类型:arxiv
- 标识:2607.11498
- 链接:https://arxiv.org/abs/2607.11498
- 主分类:multimodal
- 形态:method
- 被引:0
- 被引来源:Semantic Scholar
- S2被引:0
- 影响力被引:0
- TLDR:Pointmaps provide robot-frame 3D geometry while preserving the dense H x W grid expected by pretrained 2D VLAs, so they integrate into existing VLAs with minimal architectural change and improve both pi0.5 and SmolVLA and outperform representative camera-viewpoint and 3D-aware baselines.
- 待LLM分类:否
- 标题中文:像机器人一样看:面向视觉-语言-动作模型的机器人中心点图
- TLDR中文:Pointmaps 在保留预训练 2D VLA 所需 H × W 稠密网格的同时,提供机器人坐标系下的 3D 几何信息,能以极小的架构改动集成到现有 VLA 中,并提升 pi0.5 与 SmolVLA 的性能,优于代表性的相机视点和 3D 感知基线。
- 来源文件:
- /inbox/tom/_candidates/2026-07-21-rag-retrieval-reranking-candidates.json
- /inbox/tom/_candidates/2026-07-21-agent-rag-longcontext-candidates.json
- [S2 enrich]