本文提出 Vision-aligned Latent Reasoning(VaLR),一种简洁而有效的推理框架,在每个 Chain of Thought 推理步骤之前动态生成视觉对齐的 latent token,引导模型在 latent space 中基于感知线索进行推理。Vision-aligned Latent Reasoning (VaLR) is introduced, a simple, yet effective reasoning framework that dynamically generates vision-aligned latent tokens before each Chain of Thought reasoning step, guiding the model to reason based on perceptual cues in the latent space.
论文
313 张论文卡片 · 多模态 · 方法 · OA 绿色
本文提出 ContextRL,一种上下文感知的强化学习方法,通过间接辅助目标提升长周期推理与多模态性能,并与将相同对比上下文复用作标准 query–context–answer 样本的数据增强基线进行对比。This work proposes ContextRL, a context-aware reinforcement learning (RL) method that improves long-horizon reasoning and multimodal performance through an indirect auxiliary objective, and compares against data-augmentation baselines that repurpose the same contrastive contexts as standard query--context--answer examples.
本研究建立了一套实用的 LongPT 方案,为推进长上下文 vision-language 模型奠定了经验基础,并提出 MMProLong,无需任务专属监督即可泛化至基于网页的多模态 needle 检索、长上下文图文压缩以及长视频理解等任务。This study establishes a practical LongPT recipe and an empirical foundation for advancing long-context vision-language models, and introduces MMProLong, which generalizes to webpage-based multimodal needle retrieval, long-context vision-text compression, and long-video understanding without task-specific supervision.
本文提出 LLaDA-V,一种完全基于扩散范式的多模态大语言模型 (MLLM),将视觉指令微调与 masked diffusion 模型相结合,脱离了当前多模态方法中主流的自回归范式。LLaDA-V is introduced, a purely diffusion-based Multimodal Large Language Model (MLLM) that integrates visual instruction tuning with masked diffusion models, representing a departure from the autoregressive paradigms dominant in current multimodal approaches.
本文推出 STEP3-VL-10B,一个面向"紧凑效率与前沿级多模态智能"权衡的轻量级开源基础模型,并发布完整模型套件,为社区提供强大、高效且可复现的 baseline。STEP3-VL-10B is presented, a lightweight open-source foundation model designed to redefine the trade-off between compact efficiency and frontier-level multimodal intelligence, and the full model suite is released to provide the community with a powerful, efficient, and reproducible baseline.
本文概述了第二届 MAGMaR(Multimodal Retrieval 驱动的多模态增强生成)研讨会共享任务的成果,参赛系统聚焦于视频检索,或在给定检索视频的基础上进行有依据的文章生成。This overview paper presents the results of the shared task for the second workshop on Multimodal Augmented Generation via Multimodal Retrieval (MAGMaR), where participants submitted systems focused on either video retrieval or grounded generation of articles given retrieved videos.
我们提出 MOSS-VL,一个开源视觉-语言模型系列,将实时交互(边说边看)视为一等能力。它贯穿整个栈进行协同设计:语言解码器仅通过门控交叉注意力访问视觉,因此模型在生成过程中可以自然地感知新输入帧;合成的交互语料用于监督何时说话、何时沉默、何时修正;分阶段课程将所有实时相关训练集中在一个轻量的最终阶段,基于强大的离线基础模型。在离线场景下,MOSS-VL-Instruct 在We present MOSS-VL, an open vision-language model family that treats real-time interaction -- perceiving while it speaks -- as a first-class capability. It is co-designed across the stack: the language decoder attends to vision only through gated cross-attention, so the model can naturally see incoming frames while generating; a synthesized interaction corpus supervises when to speak, when to stay silent, and when to revise; and a staged curriculum concentrates all real-time-specific training in one light final stage over a strong offline foundation. Offline, MOSS-VL-Instruct is competitive at
性能评估显示,新训练的深度 CNN 模型 SentiBank 2.0(即 DeepSentiBank)在标注准确率与检索性能上较此前主要采用二分类 SVM 的版本有显著提升。Performance evaluation shows the newly trained deep CNNs model SentiBank 2.0 (or called DeepSentiBank) is significantly improved in both annotation accuracy and retrieval performance, compared to its predecessors which mainly use binary SVM classification models.
提出 Internalized Visual Thinking (IVT),一种在无标签视频上联合优化文本预测与下一 embedding 预测的后训练框架,表明在推理时显式的像素级生成对有效的主动视频推理并非必要。Internalized Visual Thinking (IVT), a post-training framework that jointly optimizes textual prediction and next-embedding prediction over unlabeled videos, is introduced, suggesting that explicit pixel-space generation at inference time may not be necessary for effective proactive video reasoning.
提出 EditBridge,一个用于高效超高分辨率编辑的扩散桥框架,可在最高 4K 分辨率下实现高保真编辑与卓越感知质量,在 2K 分辨率下带来 3.6–8.4 倍加速,并能在 61 秒内完成实用 4K 编辑。This work proposes EditBridge, a diffusion bridge framework for efficient ultra high-resolution editing that achieves high-fidelity editing with superior perceptual quality at resolutions up to 4K, delivering 3.6--8.4$\times$ speedup at 2K and enabling practical 4K editing in 61 seconds.
提出 DiSCO,一种零样本、严格黑盒的防御方法,完全在提示层面以即插即用模块的形式运行,无需模型重训练、微调或访问模型内部,可直接应用于任何文本到图像系统,无需对模型本身进行任何修改。DiSCO is proposed, a zero-shot, strictly black-box defense that operates entirely at the prompt level as a plug-and-play module, requiring no model retraining, fine-tuning, or access to model internals, and can be readily applied to any text-to-image system without necessitating any changes to the model itself.
本文系统性地研究了视觉编码器扩展中的 MoE 设计,发现细粒度 MoE 拓扑相较于稠密与标准 MoE 基线均带来显著提升;提出了一种无辅助损失的均衡变体以改善专家利用率,并设计了专用 MoE kernel 以缓解推理时延开销。This work systematically study MoE designs for vision encoder scaling and finds that fine-grained MoE topologies yield substantial gains over both dense and standard MoE counterparts, and proposes an auxiliary-loss-free balancing variant for better expert utilization, and designs a specialized MoE kernel to mitigate inference latency overhead.
本文提出 OmniScientist,一种端到端、全模态 AI 科学家,可直接基于异构原始证据开展跨学科研究,表明全生命周期感知对于基于证据的科学发现至关重要,并为构建广泛适用的 AI 科学家提供了一条切实可行的路径OmniScientist is introduced, an end-to-end, omni-modal AI scientist that conducts multidisciplinary research directly from heterogeneous raw evidence and demonstrates that lifecycle-wide perception is essential for evidence-grounded scientific discovery and provides a practical path toward broadly capable AI scientists.
提出 VA-Judger,一种链式思考的通用奖励模型:从质量差距明显的样本对中学习以建立结构化输出和粗粒度偏好判别,再通过拒绝采样(以人类标注验证)蒸馏出可靠的偏好解释用于更难的质量相近样本比较,最后执行维度级强化学习,将人类反馈分解到各独立质量维度以获得更稠密的奖励信号。VA-Judger is proposed, a chain-of-thought omni-reward model that learns from pairs with clear quality gaps to establish structured output and coarse preference discrimination, then distills reliable preference explanations for harder near-quality comparisons via rejection sampling verified against human annotations, and finally performs dimension-wise reinforcement learning that decomposes human feedback into individual quality dimensions for denser reward signals.
提出 RapidLiDAR,一种将初始化本身视为可学习、数据驱动组件的 LiDAR 场景补全方法,在与 SOTA 相当的补全性能下,0.1 秒完成整个场景,比此前最快方法快 2.3 倍。RapidLiDAR is presented, a LiDAR scene completion method that treats the initialization itself as a learned, data-driven component and achieves completion performance on par with the state of the art while completing a full scene in 0.1 seconds, which is 2.3 times faster than the fastest prior method.
4DAnyone 在新视角视频质量和下游 4DGS 重建上均优于先前方法,并具有稳健的野外泛化能力。4DAnyone outperforms prior methods in both novel-view video quality and downstream 4DGS reconstruction, with robust in-the-wild generalization.
在领域内和分布偏移设置下,增加测试时计算可显著提升下一子任务预测准确率,这些增益进一步转化为长时序机器人操作任务中更高的闭环成功率。Across in-domain and distribution-shifted settings, allocating additional test-time computation substantially improves next-subtask prediction accuracy, and these gains translate into higher closed-loop success on long-horizon robot manipulation tasks.
本文提出 SparsePR,一种无需训练的方法,将响应耦合划分(Response-Coupled Partitioning)与探针拟合残差重建(Probe-Fitted Residual Reconstruction)相结合,并发现划分方式既影响同一分组内查询偏好支持集之间的重叠程度,也影响稀疏输出的仿射函数对稠密与稀疏输出差异的拟合能力。This work introduces SparsePR, a training-free method combining Response-Coupled Partitioning with Probe-Fitted Residual Reconstruction, and finds that partition choice affects both the overlap among grouped queries' preferred supports and how well an affine function of the sparse output can represent the difference between dense and sparse outputs.
提出 Patch Reparameterization,在保留原始语义通路的同时,添加一个面向重构的 patch embedding,为同一组冻结的 ViT 块提供细粒度视觉信息,在保持多模态理解能力的同时实现高保真图像重构,并取得有利的重构—生成权衡。Patch Reparameterization is introduced, which preserves the original semantic pathway while adding a reconstruction-aware patch embedding that provides fine-grained visual information to the same frozen ViT blocks, and preserves multimodal understanding while enabling high-fidelity image reconstruction and a favorable reconstruction--generation trade-off.
本文提出 Block3D,一种块级扩散框架,将离散形状 token 序列划分为连续块,自回归地生成各块,并联合去噪当前块内的所有 token,同时引入置信度引导的块内修正机制,在每块定稿前对低置信度 token 进行修订。Block3D is proposed, a block-wise diffusion framework that partitions the discrete shape-token sequence into contiguous blocks, generates the blocks autoregressively, and jointly denoises all tokens within the current block and introduces confidence-guided intra-block correction, which revises low-confidence tokens before each block is finalized.
本文建立了一个包含超过 1,000 个细粒度编辑概念的综合性层次化分类体系,并提出一种密集监督训练策略,将多个互不干扰的概念合成到单个图像对中,显著提升了训练效率和模型整体性能。A comprehensive hierarchical taxonomy featuring over 1,000 fine-grained edit concepts is established and a dense supervision training strategy that synthesizes multiple non-interfering concepts into single image pairs is proposed that significantly enhances both training efficiency and overall model performance.
覆盖 4 个 FR backbone 与 2 个 VLM 编码器;EXPL-FR 无需访问模型架构,支持身份级、单图及差异式解释,并在三种监督设置(人工标注、VLM 伪标签、完全 prompt 驱动的审计)下针对真实核验行为进行属性级审计基准测试。This work covers four FR backbones and two VLM encoders, EXPL-FR needs no architecture access, and supports identity-level, per-image, and differential explanations, and benchmark attribute-level auditing under three supervision settings, human labels, VLM pseudo-labels, and the authors' fully prompt-driven audit, against real verification behavior.
介绍 Entropy-Valley (EV):一种无需训练的画布长度选择器,通过 all-mask 前向的预测平均熵对候选目标画布打分,并挑选出 backbone 最「准备好」填充的画布。Entropy-Valley (EV), a training-free length selector that scores candidate target canvases by mean predictive entropy from all-mask forward passes and selects the canvas the backbone is most prepared to fill, is introduced.
本文研究视频 MLLM 的 RL 后训练样本效率与可扩展性,并提出 OraRL——一种随模型规模与数据规模共同 scaling 的解耦 advantage estimator,在 0.8B 到 9B backbone 上均超越其基线,并在 100k prompts 规模下超越 GRPO。The sample efficiency and scalability of RL post-training for video MLLMs and introduces OraRL, a decoupled advantage estimator that scales with model size and data, surpassing its backbone from 0.8B to 9B and GRPO up to 100k prompts.
本文提出 MoTE(Mixture of Task Experts),一种将大语言模型前馈网络转化为任务特定专家同时保持多模态 backbone 共享的 decoder 架构,并在五个 COIN 基准上使用显式任务路由进行评估。This work proposes MoTE (Mixture of Task Experts), a decoder architecture that converts large language model feed-forward networks into task-specific experts while keeping the multimodal backbone shared and evaluates it on five COIN benchmarks using explicit task routes.
本文提出 GigaBrain-0.7,一种跨多种机器人 embodiment 泛化能力显著增强的 embodied foundation model,并引入一阶段对齐训练,联合优化 vision-language 理解和多 embodiment 动作生成。This work presents GigaBrain-0.7, an embodied foundation model with substantially improved generalization across diverse robot embodiments, and introduces one-stage alignment training that jointly optimizes vision-language understanding and multi-embodiment action generation.
评估表明,RubSE 在终轮和最佳轮设置下均显著优于朴素 self-evolution,refinement 轨迹更稳定,且轨迹级性能上限更高。Evaluations demonstrate that RubSE substantially outperforms na\"ive self-evolution in final-round and best-round settings, achieving more stable refinement trajectories and a higher trajectory-level performance ceiling.
结果表明,记忆、几何控制以及 rollout-aware 训练为生成连贯故事和持续演化的交互式世界提供了实用基础。Results indicate that memory, geometric control, and rollout-aware training provide a practical foundation for generating coherent stories and continuously evolving interactive worlds.
本文提出 FHM-Video,一种基于 check-before-score 原则的、由 checklist 驱动的统一数据构建框架,在 FIRM-Video-Bench 上取得最佳综合 MAE,并在三种视频生成器的 Best-of-8 采样中始终获得最高的 VBench Total、Quality 和 Semantic Score。FHM-Video, a unified checklist-driven data construction framework based on a check-before-score principle, is introduced, which achieves the best overall MAE on FIRM-Video-Bench while consistently delivering the highest VBench Total, Quality, and Semantic Scores in Best-of-8 sampling across three video generators.
本文发布一个土耳其语多模态对话数据集,包含未脚本化的双人交互,并提供同步的前向视频、可归属到每位说话人的独立音频通道以及时间对齐的转写文本。A multimodal Turkish conversational dataset of unscripted dyadic interactions, comprising synchronized front-facing video, per-speaker audio channels that allow overlapping speech to be attributed to individual speakers, and time-aligned transcriptions is introduced.
本文用前馈 4D 重建奖励替代静态 critic,显式建模场景动态,使连贯运动获得高一致性奖励,并加入对自然 scene-flow 幅值进行奖励同时抑制抖动与非刚性伪影的运动先验。This work replaces the static critic with a feed-forward 4D reconstruction reward that explicitly models scene dynamics, allowing coherent motion to receive high consistency rewards, and adds a motion prior that rewards natural scene-flow magnitude while penalizing jitter and non-rigid artifacts.
本文发布 LibriBrain100,一个面向语音解码的大规模 MEG 数据集,从设计上保证可复现、标准化评估,并展示了广泛多被试数据的价值:对预训练模型进行有监督微调可大幅弥补单被试数据不足。LibriBrain100 is introduced, a large-scale MEG dataset for speech decoding designed from the ground up for reproducible, standardised evaluation and the value of broad multi-subject data is demonstrated: supervised finetuning of a pre-trained model can substantially compensate for limited per-subject data.
本文提出一个 Agentic 视频编辑框架,利用 LLM 与 VLM 的协同实现 shot 级视频解耦与精确指令解析,并构建 MMLVE-Bench,一个聚焦 MMLVE 的数据集,具有复杂的真实世界时空动态、高密度异构指令以及稀疏随机的实体分布。This work introduces an agentic editing framework that leverages the synergy of Large Language Models (LLMs) and Vision-Language Models (VLMs) to achieve shot-level video decoupling and precise instruction parsing, and constructs MMLVE-Bench, which is an MMLVE-focused dataset characterized by complex real-world spatiotemporal dynamics, high-density heterogeneous instructions, and sparse, random entity distributions.
结果将图像编辑定位为一种专门的视觉工作空间而非通用推理机制,并将 Aphanta 确立为可复用的协议,用于度量任务-表征对齐、编辑器实现及下游 pipeline 实用性。The results position image editing as a specialized visual workspace rather than a universal reasoning mechanism, and establish Aphanta as a reusable protocol for measuring task--representation alignment, editor realization, and downstream pipeline utility.
本文提出 GameWAM,据其所知是首个面向原生闭环游戏与 GUI 控制的 WAM,并发现 Low-Frequency Action Source Imprinting (LASI):在固定条件下,采样动作源的低频分量系统性地引导生成的粗粒度相机运动,揭示了生成式控制中的源敏感性失效模式。This work introduces GameWAM, to its knowledge the first WAM for native closed-loop gameplay and GUI control and uncovers Low-Frequency Action Source Imprinting (LASI), in which low-frequency components of the sampled action source systematically steer coarse generated camera motion under fixed conditioning, revealing a source-sensitivity failure mode in generative control.
传统视频编辑主要关注场景级内容,而直播更强调人物主体。然而,直接将现有视频编辑方法应用于以人为中心的直播仍具挑战,因为它们可能引入面部表情不一致,且通常依赖多个离线推理步骤,难以满足实时交互需求。我们提出 EditaLive,一个用于实时流式角色视频编辑的新型框架。具体而言,我们从预训练图像动画模型(Wan-Animate)出发,该模型天然解耦...Conventional video editing primarily focuses on scene-level content, whereas live streaming places greater emphasis on the human subject. However, directly applying existing video-editing methods to human-centric live streaming remains challenging, as they may introduce facial-expression inconsistencies and typically depend on multiple offline inference steps, making them unsuitable for real-time interaction. We propose EditaLive, a novel framework for real-time streaming character video editing. In detail, we start from a pretrained image animation model (Wan-Animate), which naturally decoupl