本文对 20 余个近期 MLLM 进行大规模评估,结果表明视频指令跟随对当前模型仍具挑战,尤其是在涉及多约束、语义约束或需要根据视频内容选择正确分支/路径的复杂条件结构时。This work conducts a large-scale evaluation of more than 20 recent MLLMs and shows that video instruction following remains challenging for current models, especially for instructions with many constraints, semantic constraints, or complex conditional structures that require selecting the correct branch or path based on video content.
论文
443 张论文卡片 · 多模态
评估表明,RubSE 在终轮和最佳轮设置下均显著优于朴素 self-evolution,refinement 轨迹更稳定,且轨迹级性能上限更高。Evaluations demonstrate that RubSE substantially outperforms na\"ive self-evolution in final-round and best-round settings, achieving more stable refinement trajectories and a higher trajectory-level performance ceiling.
结果表明,记忆、几何控制以及 rollout-aware 训练为生成连贯故事和持续演化的交互式世界提供了实用基础。Results indicate that memory, geometric control, and rollout-aware training provide a practical foundation for generating coherent stories and continuously evolving interactive worlds.
本文提出 FHM-Video,一种基于 check-before-score 原则的、由 checklist 驱动的统一数据构建框架,在 FIRM-Video-Bench 上取得最佳综合 MAE,并在三种视频生成器的 Best-of-8 采样中始终获得最高的 VBench Total、Quality 和 Semantic Score。FHM-Video, a unified checklist-driven data construction framework based on a check-before-score principle, is introduced, which achieves the best overall MAE on FIRM-Video-Bench while consistently delivering the highest VBench Total, Quality, and Semantic Scores in Best-of-8 sampling across three video generators.
本文发布一个土耳其语多模态对话数据集,包含未脚本化的双人交互,并提供同步的前向视频、可归属到每位说话人的独立音频通道以及时间对齐的转写文本。A multimodal Turkish conversational dataset of unscripted dyadic interactions, comprising synchronized front-facing video, per-speaker audio channels that allow overlapping speech to be attributed to individual speakers, and time-aligned transcriptions is introduced.
本文用前馈 4D 重建奖励替代静态 critic,显式建模场景动态,使连贯运动获得高一致性奖励,并加入对自然 scene-flow 幅值进行奖励同时抑制抖动与非刚性伪影的运动先验。This work replaces the static critic with a feed-forward 4D reconstruction reward that explicitly models scene dynamics, allowing coherent motion to receive high consistency rewards, and adds a motion prior that rewards natural scene-flow magnitude while penalizing jitter and non-rigid artifacts.
本文发布 LibriBrain100,一个面向语音解码的大规模 MEG 数据集,从设计上保证可复现、标准化评估,并展示了广泛多被试数据的价值:对预训练模型进行有监督微调可大幅弥补单被试数据不足。LibriBrain100 is introduced, a large-scale MEG dataset for speech decoding designed from the ground up for reproducible, standardised evaluation and the value of broad multi-subject data is demonstrated: supervised finetuning of a pre-trained model can substantially compensate for limited per-subject data.
计算机使用 Agent 将自然语言指令 grounding 到截图中以定位界面元素,但现有基准无法隔离模型是否将关系语言正确绑定到对应元素。我们提出 GUI-Primitives,一个包含 994 个条目的基准,由跨七种空间关系(左右、上下、包含、对齐、邻近、列表序数、遮挡)的对比指令对组成。每对保持截图和锚点不变,仅改变关系表达,使正确目标在两个指定候选之间切换。五位标注者……Computer-use agents ground natural-language instructions in screenshots to locate interface elements, yet existing benchmarks do not isolate whether models bind relational language to the correct element. We introduce GUI-Primitives, a 994-item benchmark of contrastive instruction pairs over seven spatial relations in graphical user interfaces (left/right, above/below, containment, alignment, proximity, list ordinal, occlusion). Each pair holds the screenshot and anchor fixed while changing the relation expression, so the correct target moves between two designated candidates. Five annotators
本文提出一个 Agentic 视频编辑框架,利用 LLM 与 VLM 的协同实现 shot 级视频解耦与精确指令解析,并构建 MMLVE-Bench,一个聚焦 MMLVE 的数据集,具有复杂的真实世界时空动态、高密度异构指令以及稀疏随机的实体分布。This work introduces an agentic editing framework that leverages the synergy of Large Language Models (LLMs) and Vision-Language Models (VLMs) to achieve shot-level video decoupling and precise instruction parsing, and constructs MMLVE-Bench, which is an MMLVE-focused dataset characterized by complex real-world spatiotemporal dynamics, high-density heterogeneous instructions, and sparse, random entity distributions.
结果将图像编辑定位为一种专门的视觉工作空间而非通用推理机制,并将 Aphanta 确立为可复用的协议,用于度量任务-表征对齐、编辑器实现及下游 pipeline 实用性。The results position image editing as a specialized visual workspace rather than a universal reasoning mechanism, and establish Aphanta as a reusable protocol for measuring task--representation alignment, editor realization, and downstream pipeline utility.
本文提出 GameWAM,据其所知是首个面向原生闭环游戏与 GUI 控制的 WAM,并发现 Low-Frequency Action Source Imprinting (LASI):在固定条件下,采样动作源的低频分量系统性地引导生成的粗粒度相机运动,揭示了生成式控制中的源敏感性失效模式。This work introduces GameWAM, to its knowledge the first WAM for native closed-loop gameplay and GUI control and uncovers Low-Frequency Action Source Imprinting (LASI), in which low-frequency components of the sampled action source systematically steer coarse generated camera motion under fixed conditioning, revealing a source-sensitivity failure mode in generative control.
传统视频编辑主要关注场景级内容,而直播更强调人物主体。然而,直接将现有视频编辑方法应用于以人为中心的直播仍具挑战,因为它们可能引入面部表情不一致,且通常依赖多个离线推理步骤,难以满足实时交互需求。我们提出 EditaLive,一个用于实时流式角色视频编辑的新型框架。具体而言,我们从预训练图像动画模型(Wan-Animate)出发,该模型天然解耦...Conventional video editing primarily focuses on scene-level content, whereas live streaming places greater emphasis on the human subject. However, directly applying existing video-editing methods to human-centric live streaming remains challenging, as they may introduce facial-expression inconsistencies and typically depend on multiple offline inference steps, making them unsuitable for real-time interaction. We propose EditaLive, a novel framework for real-time streaming character video editing. In detail, we start from a pretrained image animation model (Wan-Animate), which naturally decoupl
提出 TacForcing,一种融合执行期触觉反馈的流式动作生成框架:按顺序生成动作块并保留未完成块的中间状态,并引入执行感知触觉注意力(EATA),将每次触觉更新的直接访问限制在下一个即将执行的块。TacForcing is introduced, a streaming action-generation framework incorporating execution-time tactile feedback that generates action blocks sequentially while preserving intermediate states of unfinished blocks and introduces Execution-Aware Tactile Attention (EATA), which restricts direct access to each tactile update to the next block scheduled for execution.
提出 Luce,一种三维表示方法,将几何与 PBR 材质统一到体素化的多模态高斯云中,分别使用专用高斯基元表示反照率、金属度-粗糙度与法线,在单图到三维生成任务上达到 SOTA。Luce is proposed, a 3D representation that unifies geometry and PBR materials within a voxelized multimodal Gaussian cloud, using dedicated Gaussian primitives for albedo, metallic-roughness, and surface normals to achieve state-of-the-art single-image-to-3D generation.
提出 LayerRecall,一种由当前状态条件化、按层选择的 memory router,仅从历史 K/V state 中检索相关状态并注入到 backbone 特定的 memory-sensitive 层中,其余层保留局部 attention;并提出 Cross-Horizon Prediction Matching (CHPM),借助特权的长上下文参考在预测空间中对有界 memory router 进行监督。This work introduces LayerRecall, a current-conditioned, layer-selective memory router that retrieves relevant historical K/V states and injects them only into backbone-specific memory-sensitive layers while preserving local attention elsewhere, and proposes Cross-Horizon Prediction Matching (CHPM), which uses a privileged long-context reference to supervise the bounded-memory router in prediction space.
本文提出 Judge co-adaptation from Zero data (J-Zero),一个统一的 Challenger–Solver–Judge 自进化框架,支持在可验证与不可验证领域的自我提升,并识别出 Judge 协同进化是这一持续改进的关键驱动力。This work proposes Judge co-adaptation from Zero data (J-Zero), a unified Challenger--Solver--Judge self-evolution framework that supports self-improvement across both verifiable and unverifiable domains and identifies Judge co-adaptation as the key driver of this sustained improvement.
提出 Ring Forcing,一个自回归视频扩散框架,旨在稳健构建并精确利用长期记忆,并通过稀疏 RoPE 机制实现灵活、可扩展的记忆适配,同时充分利用预训练先验。Ring Forcing is presented, an autoregressive video diffusion framework designed to robustly construct and precisely utilize long-term memory and a sparse RoPE mechanism to enable flexible, scalable memory adaptation while fully exploiting pre-trained priors.
提出 Decoupled Block Attention,在保留共享 video-query 上下文访问的同时消除跨 box 依赖,并结合用于时间边界与空间几何的 localization-aware policy optimization;并行 tube generation 被证明是视频中 autoregressive 定位的一种高效替代方案。Decoupled Block Attention is introduced, which preserves access to shared video-query context while eliminating cross-box dependencies, together with localization-aware policy optimization for temporal boundaries and spatial geometry, and parallel tube generation is shown to be an efficient and effective alternative to autoregressive localization in videos.
提出 Intention Distillation (INDI),将行为级意图蒸馏到动作解码器中,并以目标依赖的方式组织下游预测;研究表明,动作解码器显式建模其生成行为的语义目标能够带来收益。Intention Distillation (INDI) is proposed, which distills behavior-level intent into the action decoder and organizes downstream predictions in an objective-dependent manner, and shows that action decoders benefit from explicitly modeling the semantic objective of the behavior they generate.
ViT-5 的设计与当代基础模型实践保持一致,可作为对 vanilla ViT 的直接替换升级方案,适用于 2020 年代中期的视觉骨干网络,并为生成建模提供更强大的骨干。With a design aligned with contemporary foundation-model practices, ViT-5 offers a simple drop-in upgrade over vanilla ViT for mid-2020s vision backbones and serves as a stronger backbone for generative modeling.
通过发布紧凑的 7B 生成器和 2K Refiner,该工作致力于使原生音视频生成平民化,并为统一音视频生成建模的未来研究提供可及的基础。By releasing the compact 7B generator and 2K Refiner, this work seeks to democratize native audio-video generation and provide an accessible foundation for future research in unified audio-video generative modeling.
提出 MTPaperBananaBench,一个面向多轮图表生成的基准,包含 292 张图像和 3,518 条用户需求标注,并引入 PaperBanana-Interact,一个通过内部 critique-and-refine 循环来优化图表的多智能体系统。MTPaperBananaBench, a benchmark for multi-turn diagram generation containing 292 images annotated with 3,518 user requirements, is presented and PaperBanana-Interact, a multi-agent system that refines diagrams via an internal critique-and-refine loop is introduced.
提出 Hash-Atlas 网络,将 3D 场景编辑重新表述为对 2D atlas 图像的操作,从而实现 2D 编辑与 3D 重建流程的解耦。The Hash-Atlas network is proposed, which reformulates 3D scene editing as operations on 2D atlas images, thereby achieving a workflow decoupling of the 2D editing and 3D reconstruction processes.
本文利用 LSTM 网络学习视频序列表示,并通过在 UCF-101 与 HMDB-51 数据集上对人体动作识别这一监督学习任务进行微调来评估所学表示。This work uses Long Short Term Memory networks to learn representations of video sequences and evaluates the representations by finetuning them for a supervised learning problem - human action recognition on the UCF-101 and HMDB-51 datasets.
提出 SpanCalib-VLM,这是一种用于 SHROOM-Visions 共享任务的混合双系统,结合了多模态序列标注器(由 XLM-RoBERTa-Large 与 SigLIP 视觉编码器通过交叉注意力融合而成)与微调后的生成式 VLM(Qwen3.5-4B-SHROOM-SFT)。SpanCalib-VLM is presented, a hybrid dual-system for the SHROOM-Visions Shared Task that combines a multimodal sequence tagger, consisting of XLM-RoBERTa-Large fused with a SigLIP vision encoder via cross-attention, with the fine-tuned generative VLM (Qwen3.5-4B-SHROOM-SFT).
基于对数据的深入定性分析和既有理论研究,构建了一套新颖且全面的多模态错误信息分类法,并由此获得了关于社交媒体用户如何在实际场景中将图像与文本结合以传播错误信息的此前未被记录的洞见。A novel, comprehensive taxonomy of multimodal misinformation grounded in an in-depth qualitative analysis of the data and prior theoretical work is developed, which leads to previously undocumented insights about how social media users combine images with text to spread misinformation in the wild.
本文提出 RECAP-Forcing,一种无需训练的推理方法,不增加任何可学习参数,在多个强基线上稳定提升视觉质量与语义保真度,并优于现有记忆方法。This work proposes RECAP-Forcing, a training-free inference method with no additional learnable parameters that consistently improves visual quality and semantic fidelity across multiple strong baselines and outperforms existing memory methods.
本文提出 ZimaBlue,一个可扩展的框架,用于从大规模视频中学习可泛化的 World Action Models (WAMs),并采用异步 Slow-Fast 双系统架构,使生成式 WAM 具备面向实时控制的实用性。This work introduces ZimaBlue, a scalable framework for learning generalizable World Action Models (WAMs) from large-scale video, and adopts an asynchronous Slow-Fast dual-system architecture to make generative WAMs practical for real-time control.
实验表明,该方法在大幅保留通用视觉-语言能力的同时,具备出色的 3D 感知与驾驶场景理解能力;在开环、伪闭环与闭环设定下的综合评估进一步显示其运动规划性能具有很强的竞争力。Experiments demonstrate strong 3D perception and driving scene understanding while largely preserving general vision-language capability and comprehensive evaluations across open-loop, pseudo-closed-loop, and closed-loop settings further show highly competitive motion-planning performance.
本文提出 credit-addressable reasoning:推理时暴露的语义单元同时定义学习阶段比较候选与分配 credit 的位置;并实例化为 Code-CoT,保留图示、将视觉关系表示为行可寻址的可执行代码,并将推理组织为类型化事件。This work introduces credit-addressable reasoning, in which the semantic units exposed during inference also define where learning compares alternatives and assigns credit, and instantiates Code-CoT, which retains the diagram, represents visual relations as line-addressable executable code, and organizes reasoning into typed events.
VibeVoice-ASR-Streaming 是首批基于 LLM 的端到端流式说话人归属 ASR 方法之一,无需独立 diarization 阶段即可在语音到达时输出"谁说了什么"。VibeVoice-ASR-Streaming is one of the first LLM-based end-to-end approaches to streaming speaker-attributed ASR, allowing the model to produce''who said what''as speech arrives, without a separate diarization stage.
本文提出 NeoMME,一个 260M 与 800M 参数的多模态多语言双向编码器系列,可在单个双向 Transformer encoder 中处理多语言文本与原始图像 patch。This work introduces NeoMME, a family of 260M and 800M-parameter Multimodal and Multilingual bidirectional Encoders that process multilingual text and raw image patches in a single bidirectional Transformer encoder.
结果表明:领域驱动的合成数据一致优于零样本提示与朴素合成基线;主动选择相比随机采样提升了数据效率;元数据与协同过滤信号各自提升选择质量;在低资源场景下,合成数据可优于稀缺的真实对话,并进一步对真实对话形成补充。The results show that domain-grounded synthetic data consistently outperforms zero-shot prompting and naive synthetic baselines; active selection improves data efficiency over random sampling; metadata and collaborative filtering signals each improve selection quality; and, in low-resource settings, synthetic data can outperform scarce real dialogues while further complementing them.
提出 Puffin-World,一种统一的多模态架构,集成物理理解、空间仿真与 3D 世界生成重建,无需依赖外部离线模块,可支持需要多任务协同的交错式闭环应用。Puffin-World is proposed, a unified multimodal architecture that integrates physical understanding, spatial simulation, and 3D world generation and reconstruction without relying on external offline modules and enables interleaved closed-loop applications requiring synergy across multiple tasks.
提出 Temporal Context Routing,将脚本时序映射到视频与音频生成的共享时间轴上,并将每个 prompt 的引导路由到两种模态中的对应位置,同时保持与 baseline 相当的视觉质量与音视频同步性。Temporal Context Routing is introduced, which maps the script timing onto the shared temporal axis of video and audio generation and routes each prompt's guidance to the corresponding positions in both modalities, while maintaining visual quality and audio-visual synchronization comparable to those of the baselines.
过程中发现 AKS baseline 自身存在实现 bug,且两个 harness 在相同预算下运行相同已发布规则存在 0.74 分的差距,这表明此类对比应在同一受控 harness 内进行,而非跨论文比较。Along the way, an implementation bug in the own AKS baseline and a 0.74 point gap between two harnesses running the same published rules at the same budget show why these comparisons need to happen inside one controlled harness rather than across papers.