LatentPort: Beyond KV Cache - Cross-Model Transfer of Recurrent Memory in Hybrid Language Models: A 4B-to-9B Hybrid-State Handoff Without Target Prefix Replay

  • 类型:arxiv
  • 标识:2609.25053
  • 链接:https://arxiv.org/abs/2609.25053
  • 主分类:llm-infra
  • 形态:method
  • TLDR:Can one language model hand its live memory to another without the receiver rereading the context? We demonstrate useful persistent hybrid-state transfer across one architecture-matched Qwen3.5 4B-to-9B sibling pair. To our knowledge, this is the first demonstrated cross-model handoff of persistent recurrent inference state between differently sized hybrid language models without target prefix replay. Translated attention KV alone leaves a large gap; adding the Gated DeltaNet (GDN) persistent-state package lowers teacher-forced negative log-likelihood (NLL), the average next-token log-loss, by
  • 待LLM分类:否
  • 标题中文:LatentPort:超越 KV Cache——混合语言模型中循环记忆的跨模型迁移:无需目标前缀重放的 4B 到 9B 混合状态交接
  • TLDR中文:一种语言模型能否在不重读上下文的情况下将其在线记忆交接给另一个模型?我们在架构匹配的 Qwen3.5 4B 到 9B 同源模型对上演示了有用的持久混合状态迁移。据我们所知,这是首次在不进行目标前缀重放的前提下,在不同规模的混合语言模型之间实现持久循环推理状态的跨模型交接。仅迁移注意力 KV 仍留下较大差距;在此基础上叠加 Gated DeltaNet (GDN) 持久状态包,可降低 teacher-forced 负对数似然 (NLL),即平均下一 token 对数损失,幅度达
  • 来源文件
  • /inbox/tom/_candidates/2026-09-23-agent-rag-longcontext-candidates.json