研究库 论文知识库
Papers · organized/paper_cards

论文

417 张论文卡片 · 多模态 · OA 绿色

开放获取 全部 绿色 · 1640
Multimodal Speaker Verification as a Threat to Speaker Anonymization
[标题中文] 多模态说话人验证对说话人匿名化的威胁
arXiv:2607.19636 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

该工作在多语句、多模态设定下研究 ASV,考察跨匿名语音聚合信息是否影响隐私,并发现帧级聚合得到的 EER 最低。This work investigates ASV in a multi-utterance, multimodal setting and examines whether aggregating information across anonymized speech impacts privacy, and finds that frame-level aggregation yields the lowest EERs.

Three-Body Scattering for Generative Modeling
[标题中文] 用于生成建模的三体散射
arXiv:2607.18198 多模态 方法 OA · 绿色 被引 3 · S2

这些结果将 tracked scattering 确立为通往高维 one-step generation 的路径,并给出一张设计图,将 diffusion 相关监督、Drift-like 动力学与 GAN-like 目标联系起来。These results establish tracked scattering as a route to high-dimensional one-step generation and provide a design map relating diffusion-related supervision, Drift-like dynamics, and GAN-like objectives.

O-VAD: Industrial Video Anomaly Detection through Object-Centric Tracking and Reasoning
O-VAD: 基于以对象为中心的跟踪与推理的工业视频异常检测
arXiv:2607.18142 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

该工作提出一个面向异常检测的免训练 agentic 框架,无需领域特定知识,旨在追踪被检测对象随时间变化的空间-时间动态与底层变换,然后基于逐对象的时间状态轨迹进行推理,在 grounding 帧中识别异常对象。This work introduces a training-free agentic framework for anomaly detection free of domain-specific knowledge, designed to track spatial-temporal dynamics and underlying transformations of detected objects over time, and then reason over the object-wise temporal state trajectories to identify abnormal objects in grounded frames.

Rethinking Classifier-Free Guidance in On-Policy Diffusion Distillation
重新思考 On-Policy 扩散蒸馏中的无分类器引导
arXiv:2607.24731 多模态 观点 OA · 绿色 被引 5 · S2

将 Positive--Direction Matching (PDM)——一种分支感知的 OPD 目标,分别约束正预测方向与 CFG 条件方向——引入 dense-to-sparse 视频控制;由于朴素的 guided matching 对推理 guidance 尺度极为敏感,分支感知监督可实现更鲁棒、更有效的知识迁移。Positive--Direction Matching (PDM), a branch-aware OPD objective that separately constrains the positive prediction and the CFG conditional direction, is introduced to dense-to-sparse video control, where naive guided matching is highly sensitive to inference guidance scales, while branch-aware supervision enables more robust and effective knowledge transfer.

Evidence Attribution in Visual Document Understanding without Coordinates or Region Labels
无坐标与区域标签的视觉文档理解中的证据归因
arXiv:2607.24651 多模态 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

本研究考察是否存在一条实用路径,在没有坐标界面、且无需高成本区域级监督的条件下提升归因效果,并指出了这样一条可行路径。A study investigates whether there is a practical path to improve attribution without a coordinate interface and without costly region-level supervision, and indicates a practical path to improve attribution without a coordinate interface and without costly region-level supervision.

Sol-Attn: Accelerating Video Generation Inference via On-the-Fly Attention Sparsification
Sol-Attn:通过即时注意力稀疏化加速视频生成推理
arXiv:2607.24027 多模态 方法 OA · 绿色 被引 9 · S2

本文提出无需训练的 Sol-Attn(Sparsifying online attention),在单次 online-softmax pass 中统一动态路由、稀疏计算与近似修正,在稀疏注意力中取得更好的精度–效率权衡。This paper introduces training-free Sol-Attn (Sparsifying online attention), which unifies dynamic routing, sparse computation, and approximation correction in a single online-softmax pass, achieving a better accuracy-efficiency trade-off in sparse attention.

Chamaileon: Cross-Context Binder Design with Contextualized Modeling and Mixed Sampling
Chamaileon:基于上下文建模与混合采样的跨上下文结合子设计
arXiv:2607.23518 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 Chamaileon,通过将问题建模为跨上下文结合景观(cross-context binding landscape modeling),统一多目标与多态 binder 设计,有效生成可适配多样构象景观与多目标需求的序列。Chamaileon is introduced, which unifies multi-target and multi-state binder design by formulating the problem as cross-context binding landscape modeling and effectively generates sequences adaptable to diverse conformational landscapes and multi-target requirements.

WorldDiT: A Unified Diffusion Architecture for World and Action Modeling
WorldDiT:面向世界建模与动作建模的统一扩散架构
arXiv:2607.23909 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

WorldDiT 是一种统一的 diffusion Transformer 架构,将动作生成与视觉世界建模耦合,无需大型预训练 VLM 动作主干即取得强性能,在报告全部四个 suite 的方法中,其总模型参数量与平均成功率处于已报告的 Pareto 前沿上。WorldDiT, a unified diffusion transformer architecture that couples action generation with visual world modeling and achieves strong performance without a large pretrained VLM action backbone, lies on the reported Pareto frontier for total model parameters and mean success among methods reporting all four suites.

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models
UltraViT:面向大视觉-语言模型的端侧延迟优化视觉编码器
arXiv:2607.23373 多模态 应用落地 OA · 绿色 被引 0 · S2 + OpenAlex

大量实验表明,结合端上延迟感知设计与定制化训练策略,建立了高效 LVLM 编码的新 SOTA,在端上以近 1.7 倍速度运行的同时显著优于现有以编码器为中心的基线。Extensive experiments demonstrate that the on-device latency-informed design combined with the tailored training strategy establishes a new state-of-the-art for efficient LVLM encoding, significantly outperforming existing encoder-centric baselines while operating on-device at nearly 1.7xthe speed.

Visual prompt engineering for video models
视频模型的视觉提示工程
arXiv:2607.25537 多模态 方法 OA · 绿色 被引 1 · S2

本文发现视觉提示工程(visual prompt engineering,简称 VIPE)能在多项任务上提升视频推理性能,甚至比经典的文本提示工程或 test-time scaling 更有效。It is found that visual prompt engineering, or VIPE for short, improves video reasoning performance across tasks and can be even more effective than classic text-based prompt engineering or test-time scaling.

Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model
Mage-VL:面向高效编解码器原生流式多模态的基础模型
arXiv:2607.24904 多模态 方法 OA · 绿色 被引 8 · S2

本文提出 Mage-VL,一种面向实时多模态理解与交互的高效 codec-native 流式基础模型,并构建了 AI4AI 数据流水线,涵盖面向多模态 captioning 的 prompt-code 联合优化与以 AI 驱动的性能诊断,以指导训练方案。Mage-VL is presented, an efficient codec-native streaming foundation model for real-time multimodal understanding and interaction and establishes AI4AI data pipelines encompassing prompt-code joint optimization for multimodal captioning and AI-driven performance diagnosis to guide training recipes.

Parallel Decoding Distillation for Fast Image and Video Generation
并行解码蒸馏:面向快速图像与视频生成
arXiv:2607.26004 多模态 方法 OA · 绿色 被引 9 · S2

介绍 Parallel Decoding Distillation,一种简化且可扩展的基于轨迹的蒸馏方法,用于 diffusion 和 flow matching 模型的快速推理,并显著提升生成视频的多样性。Parallel Decoding Distillation is introduced, a simplified and scalable trajectory-based distillation method for fast inference of diffusion and flow matching models and presents a significant improvement in generated video diversity.

Vision Mamba: Efficient Visual Representation Learning with Bidirectional State Space Model
Vision Mamba:基于双向状态空间模型的高效视觉表示学习
arXiv:2401.09417 多模态 方法 OA · 绿色 被引 2233 · S2

本文提出基于双向 Mamba 块(Vim)的通用视觉 backbone,通过位置嵌入标记图像序列,并利用双向 state space model 压缩视觉表征,具有成为下一代视觉基础模型 backbone 的巨大潜力。This paper proposes a new generic vision backbone with bidirectional Mamba blocks (Vim), which marks the image sequences with position embeddings and compresses the visual representation with bidirectional state space models and has great potential to be the next-generation backbone for vision foundation models.

PaLM-E: An Embodied Multimodal Language Model
PaLM-E:一种具身多模态语言模型
arXiv:2303.03378 多模态 方法 OA · 绿色 被引 3257 · S2

本文提出 embodied language model,将真实世界连续传感器模态直接融入语言模型,从而建立词语与感知之间的联系,实现真实世界中的通用推理。This work proposes embodied language models to directly incorporate real-world continuous sensor modalities into language models and thereby establish the link between words and percepts to enable general inference in the real world.

SimVLM: Simple Visual Language Model Pretraining with Weak Supervision
SimVLM:基于弱监督的简单视觉语言模型预训练
arXiv:2108.10904 多模态 方法 OA · 绿色 被引 981 · S2

本文提出极简的预训练框架 SimVLM,在广泛的判别式与生成式视觉-语言基准上显著超越既往预训练方法并取得新 SOTA,包括 VQA、NLVR2 以及图像描述任务。This work presents a minimalist pretraining framework, named SimVLM, which significantly outperforms previous pretraining methods and achieves new state-of-the-art results on a wide range of discriminative and generative vision-language benchmarks, including VQA, NLVR2, and image captioning tasks.

Human-in-the-Loop Signature Bootstrapping for UAV Hyperspectral PFM-1 Mine Detection
UAV 高光谱 PFM-1 地雷检测中的人在回路签名引导
arXiv:2607.25310 多模态 应用落地 OA · 绿色 被引 1 · S2

本文研究无人机(UAV)可见光-近红外(VNIR)高光谱图像中 PFM-1 地雷的检测,使用光谱角制图(SAM)、匹配滤波器(MF)、自适应相干估计器(ACE)和约束能量最小化(CEM)。This paper studies PFM-1 landmine detection in unmanned aerial vehicle (UAV) visible and near-infrared (VNIR) HSI using spectral angle mapper (SAM), matched filter (MF), adaptive coherence estimator (ACE), and constrained energy minimization (CEM).

CLBench-V: Evaluating Multimodal Context Learning from Grounding to Knowledge Acquisition
CLBench-V: 评估多模态上下文学习——从 grounding 到知识获取
arXiv:2607.25294 多模态 评测集 OA · 绿色 被引 1 · S2

本文介绍 CLBench-V,一个多模态上下文学习 benchmark,围绕三个维度组织任务——上下文 grounding、新信息应用与新知识学习——以解决定位上下文使用失效位置的难题。This work introduces CLBench-V, a benchmark for multimodal context learning that addresses the difficulty of localizing where context use breaks down by organizing tasks around three dimensions: context grounding, new information application, and new knowledge learning.

Understanding Is Done Early: A Depth Division of Labor in Large Language Models and Its Use for Unbounded-Context Memory
理解在前部完成:大语言模型中的深度分工及其在无界上下文记忆中的应用
arXiv:2607.28263 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本工作提出 CoMem,将拆分深度 j 显式化为可复用的上下文轴:为每个 token 写入一个深度 j 的残差,选择有界 chunk 集合,并仅恢复层 [j:L)。This work introduces CoMem, which makes split depth j an explicit reusable-context axis: write one depth-j residual per token, select a bounded chunk set, and resume only layers [j:L).

ShadowDancer: Teaching Video World Models Any Action by Learning Unified Dynamics Representations from a Video and Its Shadow
ShadowDancer: 通过从视频及其阴影中学习统一动力学表示来教视频世界模型执行任意动作
arXiv:2607.28362 多模态 方法 OA · 绿色 被引 1 · S2

ShadowDancer 引入 shadow pair,即在同一动力学下对外观做独立重采样的成对视频,并由 Shadow Library 大规模构建;一个 dynamics family 可控,当且仅当能为其构造出这样的 pair。ShadowDancer introduces shadow pairs, video pairs that replay the same dynamics under independently resampled appearance, constructed at scale by the Shadow Library, so that a dynamics family becomes controllable exactly when such pairs can be constructed for it.

See2Think: Do Multimodal Models Really Use Intermediate Visual States?
See2Think:多模态模型真的使用了中间视觉状态吗?
arXiv:2607.26769 多模态 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

对代表性闭源与开源多模态模型的评测表明,视觉推理强依赖于模型与环境,没有任何单一设置能在所有任务上持续占优。Evaluating representative proprietary and open-source multimodal models, it is found that visual reasoning is strongly model- and environment-dependent, with no single setting consistently dominating across tasks.

OmniScope: Modality-Decoupled Token Compression for Omnimodal Large Language Models
OmniScope:面向全模态大语言模型的模态解耦 token 压缩
arXiv:2607.23193 多模态 方法 OA · 绿色 被引 1 · S2

提出 OmniScope,一个无需训练的 token 压缩框架,以 query 作为跨模态共享的语义锚点,并对音频与视频分别估计相关性;由此给出 OmniLLM 推理的简单设计原则:跨模态共享 query,但不共享显著性估计。This work proposes OmniScope, a training-free token compression framework that uses the query as a shared semantic anchor while estimating relevance separately for audio and video, and suggests a simple design principle for OmniLLM inference: share the query across modalities, but not the salience estimates.

RL^2-VLA: Adaptive RL Latent Compositional Steering with Test-Time Scaling for Vision-Language-Action Models
RL^2-VLA:面向 Vision-Language-Action 模型的自适应强化学习潜在组合引导与测试时缩放
arXiv:2607.26991 多模态 观点 OA · 绿色 被引 0 · S2 + OpenAlex

提出一种自适应推理时引导框架,利用VLA Latents上的强化学习,发现推理时引导在成功与失败状态下遵循根本不同的scaling laws:动作多样性在基础VLA可能失败时最为有益,但在成功可能性高时可能不必要地扰动已准确的动作。This work introduces an adaptive inference-time steering framework that leverages Reinforcement Learning on VLA Latents, and discovers that inference-time steering follows fundamentally different scaling laws under success and failure states, revealing that action diversity is most beneficial when the base VLA is likely to fail, but can unnecessarily perturb already-accurate actions when success is likely.

Not All Tokens Deserve Equal Credit: Counterfactual Sensitivity Credit Reallocation for Long-CoT Reasoning
并非所有 token 都应获得同等贡献:面向长链思维推理的反事实敏感度贡献重分配
arXiv:2607.27888 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

研究发现privileged shifts无法给出可靠的答案对齐方向,其幅度主要反映反事实敏感性而非token级学习价值;提出Counterfactual Sensitivity Credit Reallocation (CSCR),作为GRPO的简单扩展,降低高敏感token的credit并对token级优势重新归一化,同时保留原始credit预算与verifier确定的方向。These findings show that privileged shifts fail to provide reliable answer-aligned directions, while their magnitudes primarily reflect counterfactual sensitivity rather than token-level learning value, and propose Counterfactual Sensitivity Credit Reallocation (CSCR), a simple extension of GRPO that reduces credit for highly sensitive tokens and renormalizes token-level advantages to preserve both the original credit budget and verifier-determined direction.

Motion Beyond Morphology: Bootstrapping Cross-Category Motion Transfer from Abstract Motion Representations
超越形态的运动:从抽象运动表示引导跨类别运动迁移
arXiv:2608.01628 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 Motion Beyond Morphology(超越形态的运动)这一视角,旨在跨固定结构对应迁移运动,通过两阶段框架保留在不同目标形态间仍具有意义的动力学。This work introduces Motion Beyond Morphology, a perspective that seeks to transfer motion beyond fixed structural correspondence, by preserving dynamics that remain meaningful across different target morphologies by proposing a two-stage framework.

3DZip: Spatial-Aware Feature Diversity-Guided Token Compression for 3D Question Answering
3DZip:面向 3D 问答的空间感知特征多样性引导 token 压缩
arXiv:2608.01185 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 3DZip,一种三阶段 token 压缩框架:首先采用粗粒度体素化去除点级冗余,再通过 Determinantal Point Process 基于特征空间多样性选取锚点 token,最后在空间约束下融合剩余 token 以保持几何一致性。3DZip is proposed, a three-stage token compression framework that first applies coarse voxelization to remove point-level redundancy, then selects anchor tokens based on feature-space diversity via a Determinantal Point Process, and finally merges remaining tokens under spatial constraints to preserve geometric coherence.

LeapTalk: Breaking the Latency-Quality Trade-off in Talking Head Generation
LeapTalk:打破 Talking Head 生成中的延迟-质量权衡
arXiv:2608.00079 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 LeapTalk,一种通过单次前向实现稳定且实时说话头生成、可扩展至任意长视频的新颖框架,并引入音频驱动的无分类器引导机制,在极端步数缩减下保持细粒度唇形同步。LeapTalk is proposed, a novel framework that achieves stable and real-time talking-head generation with a single forward step, scaling to arbitrarily long videos, and an audio-driven classifier-free guidance mechanism to maintain fine-grained lip synchronization under extreme step reduction.

DeepVoyager-VL: Incentivizing Vision-in-the-Loop Search for Long-Horizon Multimodal Agents
DeepVoyager-VL:激励视觉在环的搜索以应对长程多模态 Agent
arXiv:2608.01827 多模态 方法 OA · 绿色 被引 1 · S2

本文提出 DeepVoyager-VL,一种面向视觉在环搜索的长程多模态深度搜索框架,通过构建多模态事件图驱动数据合成,从而产出具有中间视觉依赖与长推理链的问题。DeepVoyager-VL is proposed, a long-horizon multimodal deep-search framework for vision-in-the-loop search that constructs a multimodal event graph to drive data synthesis, yielding problems with intermediate visual dependencies and long reasoning chains.

DreamTraj: Generating 6-DoF Object Trajectories by Reading Unrendered Video Diffusion Latents
DreamTraj:通过读取未渲染的视频扩散潜在生成 6-DoF 物体轨迹
arXiv:2608.00486 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 DreamTraj,从单张 RGB 图像和任务指令预测物体 6-DoF 轨迹,推理时无需视频、深度或 CAD 模型,是首个直接从中间视频扩散表征(而非生成像素)解码物体 6-DoF 轨迹的方法This work proposes DreamTraj, which predicts a 6-DoF object trajectory from a single RGB image and a task instruction, requiring no video, depth, or CAD model at inference, and is the first approach to directly decode object 6-DoF trajectories from intermediate video diffusion representations rather than generated pixels.

Relax Within, Balance Across: Geometry-Guided Load Balancing for Vision-Language Mixture-of-Experts
内部松弛、跨模态均衡:面向视觉-语言 Mixture-of-Experts 的几何引导负载均衡
arXiv:2608.00574 多模态 方法 OA · 绿色 被引 1 · S2

在四种拆分 backbone 上,ReBA 降低所有报告 benchmark 输入的负载,同时保持与 Std-Aux 相当的平均任务准确率,并在分辨率与分块变化下降低测试范围内的平均负载与最差物理负载Across four split backbones, ReBA lowers load on every reported benchmark input while keeping mean task accuracy comparable to Std-Aux, and lowers average load over the tested range and worst physical load under resolution and tiling shifts.

A Frozen Pixel-Space Diffusion Model Can Guide Itself with Its Own Samples
冻结的像素空间扩散模型可借助自身采样进行自我引导
arXiv:2607.29122 多模态 方法 OA · 绿色 被引 3 · S2

本文在中间层附加轻量预测头,保持 backbone 冻结,并利用中间预测与最终预测的差异作为采样时的自引导方向,训练一个能够自引导的冻结预训练像素扩散模型This work attaches a lightweight prediction head to an intermediate layer, keep the backbone frozen, and use the discrepancy between the intermediate and final predictions as a self-guidance direction during sampling to train a frozen, pretrained pixel diffusion model that can guide itself.

Loud or Silent? A Reusable Framework for Per-Modality Failure Analysis in Multimodal Clinical AI
标题 -> 标题中文:响亮还是沉默?面向多模态临床 AI 的可复用逐模态失效分析框架
arXiv:2608.01462 多模态 应用落地 OA · 绿色 被引 1 · S2

一种模型无关的模态失败框架,仅依赖部署可观察信号,返回逐样本失败分类、逐模态互补矩阵(将错误归因到模态)以及响亮 vs 静默 dropout 画像(区分可监控失败与远离决策边界未被标记的失败)A model-agnostic modality-failure framework that returns a per-example failure taxonomy, a per-modality complementarity matrix that attributes error to modalities, and a loud-vs-silent dropout profile separating monitorable failures from those that pass unflagged far from the decision boundary, using only deployment-observable signals.

Seeing or Knowing? Visual Context Sensitivity in Multimodal Large Language Models
标题 -> 标题中文:看见还是知道?多模态大语言模型中的视觉上下文敏感性
arXiv:2607.26326 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

在作者研究的粗粒度属性上,MLLM 编码了视觉证据但无法可靠控制对其的依赖For the coarse attributes the authors study, MLLMs encode the visual evidence but cannot reliably control their reliance on it, indicating that for the coarse attributes they study, MLLMs cannot reliably control their reliance on it.

Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent
Video-DeepResearch:迈向下一代多模态深度研究 Agent
arXiv:2608.03979 多模态 方法 OA · 绿色 被引 2 · S2

提出 Video-DR,采用解耦的感知-探索流水线与分阶段工具解锁,强制在 web 检索前进行充分的跨帧视觉定位,实现突破模仿学习上限的自主探索。Video-DR is introduced, featuring a decoupled perception-exploration pipeline with stage-wise tool unlocking that compels exhaustive cross-frame visual grounding prior to web retrieval, enabling autonomous exploration that breaks the imitation-learning ceiling.

MiniWorld: Democratizing the Training of Video World Models from Scratch
MiniWorld:降低视频世界模型从零训练的门槛
arXiv:2608.01127 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 MiniWorld,一个从零训练流式视频世界模型的可复现框架,采用 chunk-wise 非递减噪声调度与两阶段继续训练,提升时间建模与稳定性,将促进未来视频世界模型的研究。MiniWorld is presented, a reproducible framework for training streaming video world models from scratch that adopts a chunk-wise non-decreasing noise schedule and two-stage continued training to improve temporal modeling and stability and will facilitate future research on video world modeling.

Decoding Children's Gait Behavior
解码儿童步态行为
arXiv:2608.00371 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

为人体动作识别引入新问题域:基于标准 RGB 视频对儿童步态行为进行细粒度分析,并描述一个统一的端到端框架用于解码儿科步态的基本组成。This work introduces a new problem domain for human action recognition: the fine-grained analysis of children's gait behaviors from standard RGB video, and describes a unified end-to-end framework for decoding fundamental components of pediatric gait.

Multimodal Compact Bilinear Pooling for Visual Question Answering and Visual Grounding
用于视觉问答与视觉定位的多模态紧凑双线性池化
arXiv:1606.01847 多模态 方法 OA · 绿色 被引 1616 · S2

在视觉问答与视觉定位任务上对多模态紧凑双线性池化(MCB)进行了广泛评测,结果一致表明 MCB 优于去掉 MCB 的消融版本This work extensively evaluates Multimodal Compact Bilinear pooling (MCB) on the visual question answering and grounding tasks and consistently shows the benefit of MCB over ablations without MCB.