A Glance Is All You Need: Single-Pass Fine-Grained Image Captioning with SimLoss

  • 类型:arxiv
  • 标识:2609.00591
  • 链接:https://arxiv.org/abs/2609.00591
  • 主分类:rag
  • 形态:position
  • TLDR:An image may be worth a thousand words, but most captioning models describe it in only a few. Modern vision-language models produce fluent high-level captions, yet routinely miss the attributes, counts, textures, materials, and spatial relations that make an image visually specific. Recent multi-stage systems recover some of these details through generation, decomposition, verification, and rewriting, but they do so at the expense of substantially higher inference latency. We propose SimLoss, a reference-free embedding-space objective for single-pass fine-grained image captioning. SimLoss trai
  • 待LLM分类:否
  • 来源文件
  • /inbox/tom/_candidates/2026-09-03-agent-rag-longcontext-candidates.json