arXiv:2610.10170 · RAG 检索增强
Does Document Structure Help Dense Retrieval? A Placebo-Controlled Ablation of Four Mechanisms Across Two Corpora
文档结构是否有助于稠密检索?跨两个语料库的四种机制的安慰剂对照消融
Does Document Structure Help Dense Retrieval? A Placebo-Controlled Ablation of Four Mechanisms Across Two Corpora
- 类型:arxiv
- 标识:2610.10170
- 链接:http://arxiv.org/abs/2610.10170v1
- 主分类:rag
- 形态:method
- TLDR:Retrieval-augmented generation systems increasingly rely on document-structure treatments: structure-aligned chunking, LLM-generated chunk contexts, heading-path metadata, and hierarchical two-stage retrieval. Separate studies support each on different corpora, embedders, and metrics, and none control for a shared confound: any text prepended to a chunk perturbs its embedding. We present a mechanism-isolating ablation testing all four treatments under one protocol, matching chunk sizes across conditions and adding a semantically null placebo---heading paths that are structurally valid but shuf
- 待LLM分类:否
- 标题中文:文档结构是否有助于稠密检索?跨两个语料库的四种机制的安慰剂对照消融
- TLDR中文:检索增强生成系统日益依赖文档结构处理:结构对齐分块、LLM 生成的块上下文、标题路径元数据,以及分层两阶段检索。各自研究在不同语料、嵌入器和指标上支持各自方法,但没有一项控制了共同混杂因素:在块前添加任何文本都会扰动其嵌入。我们提出机制隔离的消融方法,在同一协议下测试全部四种处理,跨条件匹配块大小,并加入一个语义为空的安慰剂——结构上有效但被打乱
- 来源文件:
- /inbox/tom/_candidates/2026-10-09-agent-rag-longcontext-candidates.json