Logit-Contribution Scoring Identifies Non-Literal Retrieval Heads
- 类型:arxiv
- 标识:2607.01002
- 链接:https://arxiv.org/abs/2607.01002
- 主分类:rag
- 形态:method
- 被引:0
- 被引来源:Semantic Scholar + OpenAlex
- S2被引:0
- OpenAlex被引:0
- 影响力被引:0
- TLDR:Logit-Contribution Scoring (LOCOS) is introduced, a write-aware detector that scores each head by the projection of its OV-circuit output onto the answer-token unembedding direction, contrasting needle and off-needle source positions in a single forward pass.
- OpenAlex ID:W7167027414
- OpenAlex DOI:10.48550/arxiv.2607.01002
- DOI:10.48550/arxiv.2607.01002
- DOI来源:OpenAlex
- 开放获取:green
- 开放获取链接:https://doi.org/10.48550/arxiv.2607.01002
- OpenAlex更新:2026-07-19
- 待LLM分类:否
- 标题中文:Logit 贡献度评分识别非字面意义检索头
- TLDR中文:提出 Logit 贡献度评分(LOCOS),一种可感知写入的检测器,通过将每个注意力头的 OV 电路输出投影到答案 token 的去嵌入方向进行打分,在单次前向传播中对比 needle 与非 needle 源位置。
- 来源文件:
- /inbox/tom/_candidates/2026-07-06-agent-memory-tool-use-candidates.json
- /inbox/tom/_candidates/2026-07-05-agent-rag-longcontext-candidates.json
- /inbox/tom/_candidates/2026-07-04-agent-rag-longcontext-candidates.json
- [S2 enrich]
- [OpenAlex backfill]