PANORAMA: Panoptic Grounded Captioning via Mask Proposal Selection

  • 类型:arxiv
  • 标识:2609.19143
  • 链接:https://arxiv.org/abs/2609.19143
  • 主分类:multimodal
  • 形态:method
  • TLDR:Intelligent systems that act in the world require image understanding that is both comprehensive and spatially grounded. Current vision-language models (VLMs) can generate fluent and detailed image captions, but reliably associating them with image pixels remains challenging. Existing methods that combine dense captioning with pixel-level grounding often produce either incomplete descriptions or inaccurate segmentation masks. We study this problem through panoptic grounded captioning, a task that requires a VLM to describe both foreground objects and background regions while grounding each ref
  • 待LLM分类:否
  • 标题中文:PANORAMA:基于掩码提议选择的全景接地描述生成
  • TLDR中文:在世界中行动的智能系统需要既全面又具备空间定位能力的图像理解。当前 vision-language models(VLM)能够生成流畅且详细的图像描述,但可靠地将描述与图像像素关联仍然具有挑战。现有将密集描述与像素级定位相结合的方法往往产生不完整的描述或不准确的分割掩码。我们通过全景接地描述生成任务研究此问题,该任务要求 VLM 同时描述前景物体与背景区域,并对每个引用...
  • 来源文件
  • /inbox/tom/_candidates/2026-09-17-agent-rag-longcontext-candidates.json