视频选题榜 2026-07-27
- https://arxiv.org/abs/2607.21553 · SANA-Video 2.0: Hybrid Linear Attention with Attention Residuals for Efficient Video Generation · 用"门控线性 + 周期 softmax 锚点 + 块级注意力残差"的混合架构,把 5B/14B 视频 DiT 推到单 H100 跑 720p 5 秒 13 秒生成,比 Wan 2.2-A14B 快 120× · round=R1
- https://arxiv.org/abs/2607.21072 · Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text · 协议化视觉回答修复 text-output vs pixel-output 接口错配,Agentic Builder 自动协议,SpatialGen-Bench 470 样本诊断 · round=R2
- https://arxiv.org/abs/2607.21653 · Molt: A Scalable PyTorch-Native Training Framework for Agentic Reinforcement Learning · NVIDIA NeMo Labs 把 codebase 压到 ~9.2K LOC 让 AI Coding Assistant 能完整推理 · round=R3