RoPE at the End of Its Rope? Theory, Diagnosis, and Mitigation of Long-Context Failures
- 类型:arxiv
- 标识:2609.39929
- 链接:http://arxiv.org/abs/2609.39929v1
- 主分类:engineering
- 形态:position
- TLDR:Long-context failures of RoPE-based language models can arise from RoPE's intrinsic tradeoff between maintaining stable token preferences and distinguishing nearby positions. Determining which weakness to address, and how, requires a more precise characterization of RoPE's behavior in trained models across context lengths. We address a key limitation of prior theory by allowing unequal query-key scales across RoPE frequencies, which aligns well with practical empirical observations. Our theory makes both vulnerabilities measurable for individual heads and inputs, and quantifies how high-freque
- 待LLM分类:是
- 标题中文:RoPE 走到尽头了吗?长上下文失效的理论、诊断与缓解
- TLDR中文:基于 RoPE 的语言模型出现长上下文失效,根源在于 RoPE 在维持稳定 token 偏好与区分相近位置之间存在固有权衡。要判断应处理哪种弱点以及如何处理,需要更精确地刻画 RoPE 在不同上下文长度下训练后模型中的行为。我们通过允许 RoPE 各频率下 query-key 尺度不一致,弥补了先前理论的一个关键局限,使之与实际经验观测高度吻合。我们的理论使得上述脆弱性对单个注意力头与输入可测量,并量化了高频
- 来源文件:
- /inbox/tom/_candidates/2026-10-01-agent-rag-longcontext-candidates.json