Behavior-Preserving KV Cache Compression

  • 类型:arxiv
  • 标识:2610.06479
  • 链接:http://arxiv.org/abs/2610.06479v1
  • 主分类:llm-infra
  • 形态:method
  • TLDR:KV caches are a major bottleneck in long-context inference and long-form generation with large language models. Existing training-free eviction policies largely rely on proxy importance signals, such as attention mass, to decide which past tokens to retain. We argue that cache compression should instead preserve the predictive behavior of the full-cache model, retaining entries whose removal would substantially change the model's output distribution. We propose Behavior-Preserving KV Cache Compression, a training-free framework that scores candidate evictions by estimating the compressed-cache
  • 副分类:engineering
  • 待LLM分类:否
  • 来源文件:
  • /inbox/tom/_candidates/2026-10-06-agent-rag-longcontext-candidates.json