Fast Weight Attention for Continual Learning
- 类型:arxiv
- 标识:2608.27763
- 链接:https://arxiv.org/abs/2608.27763
- 主分类:engineering
- 形态:method
- TLDR:Recurrent fast-weight memories and selective state-space models compress an expanding context into a fixed-size recurrent state, making the state transition an online learning rule. We study this rule under read-after-write autoregressive semantics. For the prefix-prediction objective considered here, the local fast-memory example revealed at step t is the prefix-aligned pair (x_t,y_t)=(ϕ(k_{t-1}),v_t). The common same-step association (ϕ(k_t),v_t) remains causal, but optimizes a different internal objective. We derive normalized first-order updates for squared-error regression and negative in
- 待LLM分类:是
- 来源文件:
- /inbox/tom/_candidates/2026-08-31-agent-rag-longcontext-candidates.json