Generative Late-Interaction Embeddings For Visual Document Retrieval
- 类型:arxiv
- 标识:2609.11808
- 链接:https://arxiv.org/abs/2609.11808
- 主分类:rag
- 形态:method
- TLDR:Late-interaction retrieval is the state-of-the-art for visual document search, but it pays for its accuracy in storage. Existing compression methods retain a subset or local average of the N~1,000 vectors per page. Under aggressive storage budgets, however, these methods degrade sharply, and alternatives require retraining the encoder. Investigating this degradation across three encoders, we found two consistent properties: the vectors lie exactly on the unit sphere and concentrate near a manifold of intrinsic dimension five to six. This geometry yields two insights. First, standard k-means ce
- 待LLM分类:否
- 标题中文:面向视觉文档检索的生成式晚交互 embedding
- TLDR中文:晚交互检索是当前视觉文档搜索的 SOTA 方法,但其准确性是以存储开销为代价的。现有压缩方法仅保留每页约 N~1,000 个向量中的子集或局部均值。然而在激进的存储预算下,这些方法性能急剧下降,而替代方案又需重新训练编码器。通过在三种编码器上分析这种下降现象,我们发现两个一致特性:向量恰好位于单位球面上,并集中分布在内禀维度为 5 至 6 的流形附近。该几何特性带来两点启示。首先,标准 k-means 中心……
- 来源文件:
- /inbox/tom/_candidates/2026-09-11-agent-rag-longcontext-candidates.json