CoinRAG: Contextualized Information Nugget KV Cache Reuse for Long-Context RAG
- 类型:arxiv
- 标识:2608.07458
- 链接:https://arxiv.org/abs/2608.07458
- 主分类:rag
- 形态:method
- 被引:0
- 被引来源:Semantic Scholar
- S2被引:0
- 影响力被引:0
- TLDR:This paper optimizes the Pareto frontier under low prefill latency constraints while maximizing accuracy by proposing CoinRAG (Contextualized Information Nugget KV Cache Reuse for Long-Context RAG), which seamlessly assembles their sliced KV representations with a chunk-level context.
- 副分类:llm-infra
- 待LLM分类:否
- 标题中文:CoinRAG:面向长上下文 RAG 的上下文信息要点 KV 缓存复用
- TLDR中文:本文在低 prefill 延迟约束下优化 Pareto 前沿以最大化精度,提出 CoinRAG(Contextualized Information Nugget KV Cache Reuse for Long-Context RAG),通过 chunk 级上下文无缝拼接其切片化的 KV 表示。
- 来源文件:
- /inbox/tom/_candidates/2026-08-10-agent-rag-longcontext-candidates.json
- /inbox/tom/_candidates/2026-08-13-agent-rag-longcontext-candidates.json
- [S2 enrich]