Partition the Support, Reconstruct the Residual: Training-Free Sparse Attention for Video Generation and World Models

  • 类型:arxiv
  • 标识:2608.18484
  • 链接:https://arxiv.org/abs/2608.18484
  • 主分类:multimodal
  • 形态:method
  • TLDR:Training-free block-sparse attention can accelerate video transformers, but row-wise attention concentration does not by itself specify an executable sparse operator. Queries sharing a block route may have poorly overlapping supports, while retained attention mass alone does not determine the post-softmax error from skipped interactions. We show that partition geometry affects both pooled support and the predictability of the remaining residual from the sparse output. We introduce SparsePR, which combines Response-Coupled Partitioning with Probe-Fitted Residual Reconstruction. Sampled-query ke
  • 副分类:engineering
  • 待LLM分类:否
  • 标题中文:划分支撑集、重构残差:面向视频生成与世界模型的免训练稀疏注意力
  • TLDR中文:免训练的块稀疏注意力可加速视频 Transformer,但仅凭逐行注意力集中度本身无法确定一个可执行的稀疏算子。共享同一块路由路径的查询其支撑集可能重叠很差,而仅保留的注意力质量并不足以决定由跳过的交互所产生的 softmax 后误差。我们证明划分几何同时影响池化支撑集与从稀疏输出预测剩余残差的能力。我们提出 SparsePR,将响应耦合划分与探测拟合残差重构相结合。采样查询键
  • 来源文件
  • /inbox/tom/_candidates/2026-08-25-agent-rag-longcontext-candidates.json