Fast Weight Attention for Continual Learning

  • 类型:arxiv
  • 标识:2608.27763
  • 链接:https://arxiv.org/abs/2608.27763
  • 主分类:engineering
  • 形态:method
  • TLDR:Recurrent fast-weight memories and selective state-space models compress an expanding context into a fixed-size recurrent state, making the state transition an online learning rule. We study this rule under read-after-write autoregressive semantics. For the prefix-prediction objective considered here, the local fast-memory example revealed at step t is the prefix-aligned pair (x_t,y_t)=(ϕ(k_{t-1}),v_t). The common same-step association (ϕ(k_t),v_t) remains causal, but optimizes a different internal objective. We derive normalized first-order updates for squared-error regression and negative in
  • 待LLM分类:是
  • 来源文件
  • /inbox/tom/_candidates/2026-08-31-agent-rag-longcontext-candidates.json