研究库 论文知识库
Papers · organized/paper_cards

论文

17 张论文卡片 · 多模态 · 观点 · OA 绿色

开放获取 全部 绿色 · 1640
4. Lost at the End: Primacy Bias in Multimodal Retrieval-Augmented QA
4. 迷失在末尾:多模态检索增强问答中的首因偏差
arXiv:2606.16494 多模态 观点 OA · 绿色 被引 1 · S2

研究发现,recall@k 并非已部署 KB-VQA 的正确评价指标,且弥合差距需要 reader 侧介入;本文首次对多模态 KB-VQA 中 reader 侧位置依赖性进行了受控探查,设计了一种 gold-position 协议——在问题提示中仅改变 gold passage 所在的槽位。The first controlled probe of reader-side position dependence in multimodal KB-VQA is designed, a gold-position protocol in which only the gold passage's prompt slot varies within question, indicating that recall@k is the wrong metric for deployed KB-VQA and that the remaining headroom sits on the reader side.

Locate Anything in Videos: Rethinking Efficient Generative Spatio-Temporal Video Grounding
在视频中定位任意目标:重新思考高效生成式时空视频定位
arXiv:2608.28192 多模态 观点 OA · 绿色 被引 0 · S2 + OpenAlex

提出 Decoupled Block Attention,在保留共享 video-query 上下文访问的同时消除跨 box 依赖,并结合用于时间边界与空间几何的 localization-aware policy optimization;并行 tube generation 被证明是视频中 autoregressive 定位的一种高效替代方案。Decoupled Block Attention is introduced, which preserves access to shared video-query context while eliminating cross-box dependencies, together with localization-aware policy optimization for temporal boundaries and spatial geometry, and parallel tube generation is shown to be an efficient and effective alternative to autoregressive localization in videos.

PhysStream: Streaming Physics-Grounded Video Generation with Structured Scene Memory and Fine-Grained Motion Control
PhysStream:基于结构化场景记忆与细粒度运动控制的流式物理约束视频生成
arXiv:2609.17521 多模态 观点 OA · 绿色 被引 1 · S2

PhysStream 是一个用于物理驱动图像到视频合成的自回归模型,通过引入结构化场景记忆并支持基于稀疏速度增量信号的细粒度运动控制(编码物理量),使模型能够学习底层动力学。PhysStream is an autoregressive model for physics-grounded image-to-video synthesis that incorporates structured scene memory and supports fine-grained motion control via sparse velocity-increment signals that encode physical quantities, letting the model learn the underlying dynamics.

Register Tokens for Bounded-State Reasoning in Diffusion Language Models
用于扩散语言模型有界状态推理的 Register Tokens
arXiv:2609.16372 多模态 观点 OA · 绿色 被引 0 · S2 + OpenAlex

Register被实现为专用的固定位置 token,其连续的隐藏状态被训练用于跨生成块承载推理进度,对有界代码生成尤其有效,因为正确程序通常跨越多个块。Registers are implemented as dedicated fixed-position tokens whose continuous hidden states are trained to carry reasoning progress across generation chunks, which are especially effective for bounded code generation, where correct programs usually span several chunks.

ActionPiece: Rethinking Action Tokenization for Autoregressive Vision-Language-Action Models
ActionPiece: 重新思考自回归视觉-语言-动作模型中的动作 token 化
arXiv:2609.18487 多模态 观点 OA · 绿色 被引 1 · S2

本文提出物理秩一致性 (PRC) 来衡量 tokenization 在重建后保留局部物理距离排序的程度,并提出 ActionPiece,通过对表示学习和量化的联合监督来保留物理动作关系。This work introduces physical rank consistency (PRC) to measure how well tokenization preserves local physical distance rankings after reconstruction, and presents ActionPiece, which preserves physical action relationships through joint supervision of representation learning and quantization.

1% of Tokens Can Be Enough: On Gradient Estimation in On-Policy Distillation
1% 的 token 已足够:论 On-Policy 蒸馏中的梯度估计。
arXiv:2609.24432 多模态 观点 OA · 绿色 被引 0 · S2 + OpenAlex

基于信噪分解与候选集近似的 IER(信息效率比),可基于 IER 及其与现有效用分数的组合进行 token 选择,同时保留采样的反向 KL 训练目标。An information-efficiency ratio (IER) based on a signal-to-noise decomposition and a candidate-set approximation enables token selection based on IER and its combination with existing usefulness scores, while retaining the sampled reverse-KL training objective.

Mira-Scene: Pixel-Aligned Layouts for Generative 3D Scene
Mira-Scene: Pixel-Aligned Layouts for Generative 3D Scene
arXiv:2609.23796 多模态 观点 OA · 绿色 被引 0 · S2 + OpenAlex

提出 Mira-Scene,一种组合式 3D 场景重建框架,将稀疏姿态回归替换为密集有界对应恢复,并引入多模态扩散 Transformer 联合生成物体几何与 CCM,使用模态专精的专家流配合共享注意力与位置编码以促进几何-布局一致性。Mira-Scene is presented, a compositional 3D scene reconstruction framework that replaces sparse pose regression with dense, bounded correspondence recovery and introduces a multimodal diffusion transformer that jointly generates object geometry and CCMs, using modality-specific expert streams with shared attention and positional encoding to promote geometry-layout consistency.

Uranus: Building the Next-Generation Simulation Infrastructure for Embodied AI
Uranus:为具身 AI 构建下一代仿真基础设施
arXiv:2609.24815 多模态 观点 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 Uranus,一个以关节轨迹为条件的自回归扩散模型为核心的数据驱动机器人仿真器,为不同机器人本体与相机配置下的同步多视图生成提供统一接口。This work presents Uranus, a data-driven robot simulator built around a joint-trajectory-conditioned autoregressive diffusion model, providing a unified interface for synchronized multi-view generation across diverse robot embodiments and camera configurations.

Unmask the State: When Does State Adaptation Matter for Masked Diffusion Language Models
揭示状态之谜:Masked Diffusion Language Models 中状态适应何时重要
arXiv:2609.33355 多模态 观点 OA · 绿色 被引 0 · S2 + OpenAlex

转移层面的结果表明,state adaptation 选择性而非统一地应用时最为有效,且 adaptation 的价值取决于策略反转的频率与幅度。The transition-level results suggest that state adaptation is most useful when applied selectively rather than uniformly, and that adaptation value depends on both the frequency and magnitude of strategy reversals.

Why Can't I Open My Drawer? Mitigating Object-Driven Shortcuts in Zero-Shot Compositional Action Recognition
为什么我打不开抽屉?缓解零样本组合动作识别中的物体驱动捷径
arXiv:2601.16211 多模态 观点 OA · 绿色 被引 0 · S2 + OpenAlex

本文论证了稀疏组合监督与动宾学习的不对称性会助长物体驱动的捷径学习,并指出减少捷径诊断可提升组合泛化能力。This work argues that sparse compositional supervision and verb-object learning asymmetry can promote object-driven shortcut learning and reduces shortcut diagnostics and consequently improves compositional generalization.

From Pixels to States: Rethinking Interactive World Models as Game Engines
从像素到状态:把交互式世界模型重新思考为游戏引擎
arXiv:2607.14076 多模态 观点 OA · 绿色 被引 3 · S2

本文从玩家动作控制、游戏状态动态、状态-观测持久性与实时交互生成四个维度审视交互式游戏世界建模,并针对《Black Myth: Wukong》提出可扩展的数据引擎,采集超过 90 小时的游戏画面作为状态感知型游戏世界建模的资源。This paper examines interactive game world modeling along four dimensions: player action control, game state dynamics, state-observation persistence, and real-time interactive generation, and presents a scalable data engine for Black Myth: Wukong that collects over 90 hours of gameplay as a resource for state-aware game world modeling.

ShotPlan: Cinematic Video Generation with Learnable Planning Token
ShotPlan:基于可学习规划token的电影级视频生成
arXiv:2607.17675 多模态 观点 OA · 绿色 被引 0 · S2 + OpenAlex

提出了ShotPlan,一个基于视频扩散基础模型构建的、用于显式多镜头电影级视频生成的框架,显著优于现有的电影级视频生成方法,提供更灵活的镜头管理和更强的跨镜头一致性。ShotPlan is proposed, a framework for explicit multi-shot cinematic video generation built upon a video diffusion foundation model that significantly outperforms existing cinematic video generation methods, offering more flexible shot management and stronger inter-shot consistency.

Self-Supervised Learning of Structured Dynamics from Videos
从视频中自监督学习结构化动态
arXiv:2607.21576 多模态 观点 OA · 绿色 被引 0 · S2 + OpenAlex

提出结构化动态模型(SDM),通过未来特征预测,显式地将时间变化的主导来源与残差动态分离开来,而非使用单一纠缠的隐变量或非结构化的、空间密集的转移 token 来表示视频变化。The Structured Dynamics Model (SDM) is proposed, which explicitly separates the dominant source of temporal change from residual dynamics through future-feature prediction, rather than representing video change with a single entangled latent or with unstructured, spatially dense transition tokens.

Rethinking Classifier-Free Guidance in On-Policy Diffusion Distillation
重新思考 On-Policy 扩散蒸馏中的无分类器引导
arXiv:2607.24731 多模态 观点 OA · 绿色 被引 5 · S2

将 Positive--Direction Matching (PDM)——一种分支感知的 OPD 目标,分别约束正预测方向与 CFG 条件方向——引入 dense-to-sparse 视频控制;由于朴素的 guided matching 对推理 guidance 尺度极为敏感,分支感知监督可实现更鲁棒、更有效的知识迁移。Positive--Direction Matching (PDM), a branch-aware OPD objective that separately constrains the positive prediction and the CFG conditional direction, is introduced to dense-to-sparse video control, where naive guided matching is highly sensitive to inference guidance scales, while branch-aware supervision enables more robust and effective knowledge transfer.

RL^2-VLA: Adaptive RL Latent Compositional Steering with Test-Time Scaling for Vision-Language-Action Models
RL^2-VLA:面向 Vision-Language-Action 模型的自适应强化学习潜在组合引导与测试时缩放
arXiv:2607.26991 多模态 观点 OA · 绿色 被引 0 · S2 + OpenAlex

提出一种自适应推理时引导框架,利用VLA Latents上的强化学习,发现推理时引导在成功与失败状态下遵循根本不同的scaling laws:动作多样性在基础VLA可能失败时最为有益,但在成功可能性高时可能不必要地扰动已准确的动作。This work introduces an adaptive inference-time steering framework that leverages Reinforcement Learning on VLA Latents, and discovers that inference-time steering follows fundamentally different scaling laws under success and failure states, revealing that action diversity is most beneficial when the base VLA is likely to fail, but can unnecessarily perturb already-accurate actions when success is likely.

EffectLearner: World-Aware Object-Effect Reasoning for Real-World Video Object Removal
EffectLearner:面向真实世界视频物体移除的世界感知物体-效果推理
arXiv:2608.05565 多模态 观点 OA · 绿色 被引 0 · S2 + OpenAlex

EffectLearner 是一个语义推理增强框架,结合基于 VLM 的 Object-Effect Reasoner 与基于 DiT 的 Video Eraser,在 EffectWorld-Eval 和具有挑战性的 EffectWorld-Wild 上均取得明显优势,证明其能在复杂真实场景中实现高质量的视频物体擦除。EffectLearner is proposed, a semantic-reasoning-enhanced framework that combines a VLM-based Object-Effect Reasoner with a DiT-based Video Eraser that achieves clear advantages on both EffectWorld-Eval and the challenging EffectWorld-Wild, demonstrating its ability to deliver high-quality video object removal in complex real-world scenes.

KVAE: Family of Tokenizers for Multimodal Generative Models
KVAE:面向多模态生成模型的 tokenizer 家族
arXiv:2608.05798 多模态 观点 OA · 绿色 被引 0 · S2 + OpenAlex

在客观与主观指标上的重建和生成结果匹配或超越前沿开源 tokenizer,包括 Wan-2.2、HunyuanVideo-1.5、FLUX.2、MovieGen、StableAudio 和 MMAudio 的 VAE。It is demonstrated that reconstruction and generation results on objective and subjective metrics matches or surpasses frontier opensource tokenizers, such as VAEs from Wan-2.2, HunyuanVideo-1.5, FLUX.2, MovieGen, StableAudio and MMAudio.