MoME: Mixture-of-Memory Embeddings for Context-Aware Sparse Lookup

  • 类型:arxiv
  • 标识:2609.15126
  • 链接:https://arxiv.org/abs/2609.15126
  • 主分类:rag
  • 形态:method
  • TLDR:Scaling large language models efficiently has motivated sparse capacity mechanisms such as Mixture-of-Experts and, more recently, conditional memory: token-indexed embedding tables that augment the backbone with cheap parametric lookups. Existing memory-embedding methods retrieve via a deterministic function of the surface form, which collapses different contextual senses of the same token (e.g., python the language vs. the animal) into a single fixed entry. We introduce Mixture of Memory Embeddings (MoME), a context-aware memory mechanism that replaces each token's single memory row with a mi
  • 待LLM分类:否
  • 来源文件
  • /inbox/tom/_candidates/2026-09-21-agent-rag-longcontext-candidates.json