Is Memorization Context-Sensitive? Prefix-Based Extraction Beyond Isolated Prefixes

  • 类型:arxiv
  • 标识:2610.12085
  • 链接:http://arxiv.org/abs/2610.12085v1
  • 主分类:rag
  • 形态:application
  • TLDR:Large language models (LLMs) can expose memorized training sequences under prefix-based extraction: given a prefix from a training example, the model may assign high probability to the original continuation. In deployed systems, however, prefixes are rarely evaluated in isolation. They often appear together with instructions, retrieved documents, or other task-specific context, as in retrieval-augmented generation (RAG). This motivates examining whether contextual conditioning mitigates memorization or merely changes the set of memorized samples that become extractable. We investigate this iss
  • 副分类:engineering
  • 待LLM分类:否
  • 来源文件:
  • /inbox/tom/_candidates/2026-10-09-agent-rag-longcontext-candidates.json