介绍 Tri-PvP,一个包含 8,000 样本、跨越视觉、音频与文本的三模态冲突基准,揭示了证据形式偏倚的系统性不对称:模型在视觉上更偏向感知信号,而在音频上更偏向命题性信号。Tri-PvP, an 8,000-sample tri-modal conflict benchmark crossing vision, audio, and text, is introduced, revealing a systematic asymmetry in evidence-form bias: models exhibit a stronger bias toward perceptual signal in vision but propositional in audio.
论文
443 张论文卡片 · 多模态
介绍 RecCAR(Reciprocal Cross-modal Attention Regularization),一种 KL 正则化器,利用已确立的视频到模态对应关系作为固定参考,将较弱的模态到视频对应关系向其对齐。RecCAR, standing for Reciprocal Cross-modal Attention Regularization, is introduced, a KL regularizer that uses the well-established video-to-modality correspondence as a fixed reference and aligns the weaker modality-to-video correspondence toward it.
介绍 Spatial-Interactor,一个通过交互训练 VLM 建模物理世界状态转换的框架,并将学习过程组织为三级课程:L1 被动世界状态转换、L2 主动自我状态转换、L3 长时交互轨迹。Spatial-Interactor is introduced, a framework that trains VLMs to model physical-world state transitions through interaction, organizing this learning process into a three-level curriculum covering L1 passive world-state transitions, L2 active self-state transitions, and L3 long-horizon interaction trajectories.
本文提出 Uranus,一个以关节轨迹为条件的自回归扩散模型为核心的数据驱动机器人仿真器,为不同机器人本体与相机配置下的同步多视图生成提供统一接口。This work presents Uranus, a data-driven robot simulator built around a joint-trajectory-conditioned autoregressive diffusion model, providing a unified interface for synchronized multi-view generation across diverse robot embodiments and camera configurations.
本文提出 X-Planner,一个面向具身推理的规划前端,同时解决监督与表征问题,并描述了规划文本质量与下游执行情况。This work presents X-Planner, a planning front-end that addresses both the supervision and representation of embodied reasoning, and describes planning-text quality and downstream execution.
Knowledge Pull Requests 相比从来源重写或从零重新生成,能整合更多信息并更好地保留已有内容,同时每生成一个 token 增加的信息量最多。Knowledge Pull Requests integrate more information and better preserve existing content than rewriting from sources or regenerating from scratch, while adding the most information per token generated.
本文提出 ViRDM,一种无需教师与评论家网络的视频后训练方案,将三网络蒸馏转化为仅生成器的后训练,在降低 GPU 显存与训练时间的同时提升视频质量。ViRDM, a teacher- and critic-free video post-training recipe that turns three-network distillation into generator-only post-training, reducing GPU memory use and training time while improving video quality, is introduced.
本文提出 OmniEchoBench,一个面向空间音视频感知与音-视-语言导航的统一基准,以及一个空间感知的全模态模型,该模型在预训练语义音频通路之外引入 FOA 空间编码器。OmniEchoBench, a unified benchmark for spatial audio-visual perception and audio-vision-language navigation, and a spatially aware omni-modal model, which introduces an FOA spatial encoder alongside a pretrained semantic audio pathway.
结果显示 SAE 以分布式、类别依赖的形式定位形态句法信息,而非通过原子化的语法特征。The results show that SAEs localise morpho-syntactic information in a distributed and category-dependent form rather than through atomic grammatical features.
本文提出 DeltaWAM,通过 dense-anchor、sparse-delta 与 action 三条流联合预测视觉 delta 与动作,并设计了三种在表征与计算共享上有所不同的架构;同时开发了 Streaming Delta Memory (SDM),使用紧凑的观测 delta 更新缓存的 anchor 上下文,从而减少繁重的 video-expert 处理。This work proposes DeltaWAM, which jointly predicts visual deltas and actions using dense-anchor, sparse-delta, and action streams, with three architectures that differ in representation and computation sharing, and develops Streaming Delta Memory (SDM), which updates cached anchor context with compact observed deltas, reducing heavy video-expert processing.
本文提出 AV-GRPO,一种以模态为锚点的在线 diffusion RL 框架,以及 5DAV,一种解耦的、难度可调节的训练数据集;在 LoRA 与全量 fine-tuning 条件下,其在生成质量、语义对齐和跨模态同步性上均优于 LTX-2.3。This work proposes AV-GRPO, a modality-anchored online diffusion RL framework, and 5DAV, a decoupled, difficulty-controllable training dataset that outperforms LTX-2.3 in generation quality, semantic alignment and cross-modal synchronization under LoRA and full fine-tuning.
本文提出一种可调深度网络架构,借助适配器残差模块可在线切换至不同视觉域;同时引入 Visual Decathlon Challenge 基准,用于评估表征同时捕获十个差异显著视觉域的能力,并衡量其跨域均匀识别的能力。This paper develops a tunable deep network architecture that, by means of adapter residual modules, can be steered on the fly to diverse visual domains and introduces the Visual Decathlon Challenge, a benchmark that evaluates the ability of representations to capture simultaneously ten very differentVisual domains and measures their ability to recognize well uniformly.
TrackEverything 是首个在 40 GB GPU 显存内即可在超过 1000 帧的视频中跟踪所有可见点的 3D tracker;本文还提出 3D WAFT,用场景点云内的高效特征采样取代了显存开销巨大的 4D correlation volumes。TrackEverything is the first 3D tracker capable of tracking all visible points across videos exceeding 1000 frames within 40 GB of GPU memory, and 3D WAFT is proposed, replacing memory-prohibitive 4D correlation volumes with efficient feature sampling in the scene cloud.
本文提出一条全自动的数据生成流程,可在无需任何人工标注的情况下生成图像对之间的逐点感知距离标签,并证明 diffusion trajectory 与人类视觉系统高度一致。This paper proposes a fully automated data generation pipeline that generates pointwise perceptual distance labels between image pairs without any human annotation, and demonstrates that the diffusion trajectory aligns well with the human visual system.
本文提出一种改进的 Stable Diffusion 3 架构,采用裁剪后的文本流与逐 patch 的归一化策略,使其能在 LiDAR 数据上稳定训练,并实现从自然图像到高程图的迁移;研究表明多模态条件输入可提升高程精度。A modified Stable Diffusion 3 architecture with a pruned text stream and a patch-wise normalization strategy is introduced, enabling stable training on LiDAR data and transfer from natural images to elevation maps, demonstrating that multimodal conditioning improves elevation accuracy.
BoundInk 是一种以书写者为条件的框架,将字符间的边界视为显式生成单元,既能保留书写者特有的字形外观,又能在完整文本行内改善连接性与字距。BoundInk is introduced, a writer-conditioned framework that treats inter-character boundaries as explicit generation units and preserves writer-specific glyph appearance while improving connectivity and spacing across complete text lines.
本文提出 VLA-Precision,一个面向真实世界的高效在线 RL 框架,包含 Asymmetric Co-Bootstrapping (ACoB) 算法与 ACoB-Stream 架构;其中 ACoB-Stream 以不变状态解耦与按需流式传输为设计原则,构建了经验-策略的闭环架构,可实现最高 10.9% 的吞吐与计算效率提升。VLA-Precision, an efficient real-world online RL framework featuring the Asymmetric Co-Bootstrapping (ACoB) algorithm and the ACoB-Stream architecture, and ACoB-Stream, a closed-loop experience--policy architecture that establishes invariant-state decoupling and on-demand streaming as design principles, deliver up to 10.9% improvements in throughput and computational efficiency.
LightMIS 通过 Scale-Aligned Projection 块将五级编码器的输出对齐至统一分辨率,仅聚合一次,并利用结合 Adaptive Kernel Fusion 与所提 Progressive Receptive Fusion 模块的 Adaptive Fusion Cascade 精炼融合表示。LightMIS aligns the outputs of a five-level encoder to a common resolution using Scale-Aligned Projection blocks, aggregates them once, and refines the fused representation with an Adaptive Fusion Cascade, which combines Adaptive Kernel Fusion with the proposed Progressive Receptive Fusion module.
LastOPD 仅在最后一层状态(LM head 共同读取的接口)施加潜在信号,且仅在进入 token 级 OPD 前的 10 步交叉淡入阶段施加,从而保留潜在信号中的有用部分,并在坍缩发生前将 student 交由 token 级监督。LastOPD is proposed, which applies the latent signal only at the last-layer state, the common interface both LM heads read, and only during a 10-step crossfade into token-level OPD, which keeps the useful part of the latent signal and hands the student to token-level supervision before the collapse sets in.
InfiniHand 是一个端到端流式前馈框架,联合估计 MANO 参数、相机轨迹与手部位置,直接从未标定的自我中心视频出发,相比 ViDiHand 在 ARCTIC PA-p 上降低 21%,并显著缓解世界空间漂移。InfiniHand is an end-to-end streaming feed-forward framework that jointly estimates MANO parameters, camera trajectories, and hand locations directly from uncalibrated egocentric video, achieving a 21% reduction in ARCTIC PA-p compared to ViDiHand while substantially mitigating world-space drift.
GeoVerse 是一个在世界一致条件下合成新视角的框架,在预训练 3D 基础模型的几何潜在空间内进行生成,并通过 ControlNet 风格 adapter 注入视频生成模型的外观先验。GeoVerse is a framework that synthesizes world-consistent novel views by performing generation within the geometric latent space of a pretrained 3D foundation model and injecting appearance priors from a video generative model via a ControlNet-style adapter.
UMM-Reflection 在统一模型内利用强化学习(RL)完成完整反思轨迹:兄弟轨迹共享同一初始图像,因此组相对优势可比较不同反思策略,而单条轨迹级优势同时更新反思 token 与基于 flow 的修订,避免了逐轮信用分配的组合爆炸。UMM-Reflection is introduced, which applies reinforcement learning (RL) to complete reflection trajectories inside one unified model: sibling trajectories share one initial image, so the group-relative advantage compares reflection strategies, and one trajectory-level advantage updates both the reflection tokens and the flow-based revisions, avoiding the combinatorial blow-up of per-round credit assignment.
本文提出 FlowTool,一个通过 conditional rectified flow 直接建模以输入图像和用户指令为条件的高质量工具参数分布的框架,并显著提升了推理效率。This work introduces FlowTool, a framework that directly models the distribution of high-quality tool parameters conditioned on the input image and user instruction using conditional rectified flow, and significantly improves inference efficiency.
本文勾勒了一幅计算机视觉版图:在其中,越来越复杂的视觉任务可通过通用接口访问,而精确且对保真度敏感的感知仍是重要前沿。A changing landscape of computer vision is mapped in which increasingly sophisticated visual tasks are accessible through a general-purpose interface, while precise and fidelity-sensitive perception remains an important frontier.
OLIVE 通过从持续演化的 student 生成前缀,在离线蒸馏进入平台期后继续提升,同时更好地保持了 student 的通用能力与可塑性。By regenerating prefixes from the evolving student, OLIVE continues improving after offline distillation plateaus while better preserving the general capabilities and plasticity of the student.
提出解耦信用自蒸馏,理论上将信用方向与幅度解耦为两个可靠信号,并据此校准特权教师监督,从而完成策略优化的 step-to-token 信用分配。Decoupled Credit Self-Distillation is introduced, which theoretically decouples credit direction and magnitude into two reliable signals and uses them to calibrate privileged teacher supervision, enabling step-to-token credit assignment for policy optimization.
NVV-SuperBench 与人工听力评测的结果显示,NVAlign 在标签跟随准确率上优于 SFT 与 Flow-GRPO 基线,证明直接对奖励梯度进行优化可提升连续自回归流匹配 TTS 中的非语言控制能力。Results from NVV-SuperBench and human listening evaluations show that NVAlign improves tag-following accuracy over SFT and Flow-GRPO baselines, demonstrating that direct reward-gradient optimization can improve non-verbal control in continuous autoregressive flow-matching TTS.
构建了 RoboGenesis,一个基于仿真的工作流与数据引擎,可从原子技能组合配置好的实验工作流,对 rollout 进行验证与过滤,并跨支持的机器人配置导出结构化演示数据。RoboGenesis is built, a simulation-based workflow and data engine that composes configured laboratory workflows from atomic skills, validates and filters rollouts, and exports structured demonstrations across supported robot profiles.
结果表明,harness 层面的能力组合是构建广泛、可演进 Omni 系统的实用路径,且无需改变宿主 agent 的推理核心。The results establish harness-level capability composition as a practical route to broad, evolvable Omni systems without changing the host agent's reasoning core.
提出 Omni-Decision,一个基于证据账本规划的 omni-modal agent:用显式的证据账本替代不断膨胀的对话历史,记录仍缺失的证据、已确认的内容以及记录间的冲突。Omni-Decision is presented, an omni-modal agent built on evidence-ledger planning: it replaces the growing dialogue history with an explicit evidence ledger that records what evidence is still missing, what has been confirmed, and where records conflict.
APM-Bench 将真实场景的流式交互重构为多会话生命轨迹,并揭示了一个清晰的效用—延迟—存储权衡:现有方法仍难以同时实现可靠的长程记忆、低开销以及跨会话有效的主动协助。APM-Bench is introduced, which reformulates real-world streaming interaction as multi-session life trajectories and reveals a clear utility--latency--storage trade-off: existing methods still struggle to simultaneously achieve reliable long-term recall, low overhead, and effective proactive assistance across sessions.
一种随机少步生成器,其 FID 在采样预算从 4 增加到 50 NFE 时不会下降,且相同配方可迁移到 text-to-image 生成。A stochastic few-step generator whose FID does not degrade as the sampling budget grows from 4 to 50 NFE, and the same recipe transfers to text-to-image generation.
提出 Budgeted ATTA:测试批次中仅有一小部分可获得标签,且监督时机是 active test-time adaptation 中一个关键但尚未充分探索的方面。Budgeted ATTA is introduced in which labels are available for only a fraction of test batches, and the timing of supervision is a key, yet underexplored, aspect of active test-time adaptation.
本文提出 DashboardQA,这是首个明确设计用于评估视觉-语言 GUI Agent 对真实世界仪表板理解与交互能力的基准,结果表明交互式仪表板推理对所有受评估的 VLM 而言都是一项具有挑战性的任务。DashboardQA is introduced, the first benchmark explicitly designed to assess how vision-language GUI agents comprehend and interact with real-world dashboards, and indicates that interactive dashboard reasoning is a challenging task overall for all the VLMs evaluated.
本文提出 AutoRef,在保持两侧模型冻结的前提下自动优化 harness:一个 coding Agent 迭代重写 harness 代码,并发现 AutoRef-Harness,可改进开源权重模型 FLUX。This work proposes AutoRef, which optimizes the harness automatically while keeping both models frozen: a coding agent iteratively rewrites the harness code, and discovers AutoRef-Harness, which improves the open-weight FLUX.
在包括多视角空间推理和碰撞预测在内的六项多样化视觉推理任务中,ReImaGin 始终优于纯文本推理和专用 vision-tool 基线,提升幅度高达 25%,证明了灵活、可生成的视觉推理的优势。Across six diverse visual reasoning tasks including multi-view spatial reasoning and collision prediction, ReImaGin consistently outperforms both text-only reasoning and specialist vision-tool baselines, with gains of up to 25\%, demonstrating the advantage of flexible, generative visual reasoning.