Papers · organized/paper_cards

论文

258 张论文卡片 · 多模态

开放获取 全部 绿色 · 724
Why Can't I Open My Drawer? Mitigating Object-Driven Shortcuts in Zero-Shot Compositional Action Recognition
为什么我打不开抽屉?缓解零样本组合动作识别中的物体驱动捷径
arXiv:2601.16211 多模态 观点 OA · 绿色 被引 0 · S2 + OpenAlex

本文论证了稀疏组合监督与动宾学习的不对称性会助长物体驱动的捷径学习,并指出减少捷径诊断可提升组合泛化能力。This work argues that sparse compositional supervision and verb-object learning asymmetry can promote object-driven shortcut learning and reduces shortcut diagnostics and consequently improves compositional generalization.

LongE2V: Long-Horizon Event-based Video Reconstruction, Prediction, and Frame Interpolation with Video Diffusion Models
LongE2V:基于视频扩散模型的长时间跨度事件驱动视频重建、预测与帧插值
arXiv:2607.08770 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 LongE2V,一种利用预训练视频扩散先验来联合处理基于事件视频重建、预测与帧间插值的新方法,并引入自回归展开与自适应上下文切换机制,以缓解超长序列中的时序漂移问题。This work proposes LongE2V, a novel approach that leverages pre-trained video diffusion priors to jointly handle event-based video reconstruction, prediction, and frame interpolation, and introduces Autoregressive Unrolling and Adaptive Context Switching to mitigate temporal drift in extremely long sequences.

A Quantized Native Runtime for On-Device Semantic Audio Generation
一种用于端侧语义音频生成的量化原生运行时
arXiv:2607.08526 多模态 应用落地 OA · 绿色 被引 0 · S2 + OpenAlex

该紧凑且量化原生、带内置控制的运行时为物联网音频场景下的端侧语义音频提供了实用基础;通过对转向接口的案例分析,可生成在部分属性上具有真实但有界控制的、承载口味联想的音乐。A compact, quantized runtime with built-in control a practical basis for on-device semantic audio in Internet-of-Sounds settings and a case study of the steering interface generates music carrying taste associations with genuine but bounded control for a subset of attributes.

SAM-MT: Real-Time Interactive Multi-Target Video Segmentation
SAM-MT:实时交互式多目标视频分割
arXiv:2607.08688 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

SAM-MT 成功将延迟与目标数量解耦,在保持 SAM2 鲁棒视频分割性能的同时,实现了与单目标基线相当的实时速度。SAM-MT successfully decouples latency from the number of targets, achieving real-time speed on par with single-target baselines while maintaining SAM2's robust video segmentation performance.

Phone Segmentation and Recognition through Phonological Activation Mapping
基于音韵激活映射的音素切分与识别
arXiv:2607.09020 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

论文主张语音结构已隐含在自监督语音模型(S3M)的表征中,只需对其进行引导即可同时完成切分与识别任务。It is argued that phonetic structure is already latent in the representations of self-supervised speech models (S3Ms), and one only needs to steer them to solve both segmentation and recognition tasks.

MedPMC: A Systematic Framework for Scaling High-Fidelity Medical Multimodal Data for Foundation Models
MedPMC:面向基础模型的高保真医学多模态数据规模化系统框架
arXiv:2607.07673 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 MedPMC——一种自动化、可持续更新的框架,可将宽松许可的文献转化为面向医学多模态模型的高保真基础设施,并公开发布该框架、语料库、基准与预训练模型。MedPMC, an automated, continuously updatable framework that transforms permissively licensed literature into high-fidelity infrastructure for medical multimodal models, is introduced and publicly release the framework, corpus, benchmarks, and pretrained models.

VaseMuseum: Digital Intelligent Museum for Ancient Greek Pottery
VaseMuseum:古希腊陶器数字智能博物馆
arXiv:2607.06374 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 VaseMuseum——一种面向古希腊陶器智能数字博物馆的轻量化、模块化多模态智能体框架,相比启用搜索的 VLM 基线,它提升了引用有效性,减少了知识密集型查询中的幻觉,并在含歧义场景下给出更中立的回答。VaseMuseum is proposed, a lightweight and modular multimodal agent framework for intelligent digital museums of ancient Greek pottery that improves citation validity, reduces hallucinations on knowledge-intensive queries, and produces more neutral answers under ambiguity compared with search-enabled VLM baselines.

Motion4Motion: Motion Transfer Across Subjects at Inference
Motion4Motion:推理阶段的跨主体运动迁移
arXiv:2607.11644 多模态 应用落地 OA · 绿色 被引 0 · S2 + OpenAlex

Motion4Motion 对视频中角色的 motion flow(而非骨骼)进行建模,使跨物种运动迁移更加容易。Motion4Motionmodels the motion flow of the character in a video instead of skeletons, which makes motion transfer across species easier, which makes motion transfer across species easier.

4D Human-Scene Reconstruction from Low-Overlap Captures
低重叠度采集下的 4D 人体场景重建
arXiv:2607.09125 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 StudioRecon,一种通过解耦背景与人体、并利用视频扩散模型合成数百个相机可控新视角,从稀疏低重叠相机重建 4D 人体场景的流水线,在四个真实数据集上达到了 SOTA 的新视角合成效果StudioRecon is proposed, a pipeline that reconstructs 4D human scenes from sparse, low-overlap cameras by decoupling background and humans by synthesizing hundreds of camera-controlled novel views with a video diffusion model and achieves state-of-the-art novel view synthesis across four real-world datasets.

CtrlVTON: Controllable Virtual Try-On via Visual-Instance-Prompt Segmentation
CtrlVTON:基于视觉实例提示分割的可控虚拟试穿
arXiv:2607.09362 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

推出 CtrlVTON,一个将试穿重构为图像编辑问题并引入分割掩码作为对服装布局(包括风格、尺寸与身体空间位置)像素级控制的可控 VTO 框架CtrlVTON is introduced, a controllable VTO framework that recasts try-on as an image editing problem and adds segmentation masks as pixel-level control over garment layout, including style, size, and spatial placement on the body.

LATO.2: Factorized 3D Mesh Generation with Vertex and Topology Flow
LATO.2:基于顶点流与拓扑流分解的 3D 网格生成
arXiv:2607.10623 多模态 方法 OA · 绿色 被引 2 · S2

提出 LATO.2,一个因子化 flow matching 框架,将网格生成分解为 vertex flow 和随后以已实现顶点为条件的 connectivity flow,在几何保真度和连通性质量上超越 SOTA 的拓扑感知网格生成方法。LATO.2, a factorized flow matching framework that decomposes mesh generation into a vertex flow followed by a connectivity flow conditioned on the realized vertices, is presented, which surpasses state-of-the-art topology-aware mesh generators in geometric fidelity and connectivity quality.

Latent-Identity Tuning in Text-to-Image Personalization Models
文本到图像个性化模型中的潜空间身份调优
arXiv:2607.11885 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文探索了一个预训练、冻结 encoder 的潜空间用于 text-to-image 个性化,并表明可以在该空间及由选定 token 定义的子空间中识别出有意义的编辑方向,从而实现局部化、细粒度且语义一致的编辑。This work explores the latent space of a pre-trained, frozen encoder for text-to-image personalization, and shows that meaningful directions can be identified within this space and within subspaces defined by selected tokens, enabling localized, fine-grained, and semantically coherent edits.

Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model
Xiaomi-Robotics-U0:基于 World Foundation Model 的统一具身合成
arXiv:2607.11643 多模态 应用落地 OA · 绿色 被引 1 · S2

Xiaomi-Robotics-U0 是首个支持跨多种机器人本体的高质量多视角场景生成、并引入结构化、可控的具身迁移以实现细粒度编辑的模型,同时保持多视角一致性与交互动态。Xiaomi-Robotics-U0 is the first model to support high-quality multi-view scene generation across multiple robot embodiments and to introduce structured, controllable embodied transfer for fine-grained editing while preserving multi-view consistency and interaction dynamics.

Evidence-Backed Video Question Answering
证据支撑的视频问答
arXiv:2607.11862 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 ST-Evidence,首个同时面向判别式和生成式像素级 grounding 的人工验证 benchmark,并开发可扩展的自动生成流程,构建了 16 万规模、衔接高层推理与细粒度 grounding 的数据集 ST-Evidence-Instruct。ST-Evidence is introduced, the first human-verified benchmark for both discriminative and generative pixel-level grounding, and scalable, automated generation pipelines are developed to create ST-Evidence-Instruct, a 160k-scale dataset bridging high-level reasoning with fine-grained grounding.

A Theory of Contrastive Learning with Natural Images
自然图像对比学习的一种理论
arXiv:2607.07470 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

针对一系列基本增广与任意具有平稳统计量的图像数据集,以解析方式根据对比损失计算最优表示,结果表明对于某些增广,最优解可由第一层滤波器为正弦函数的 CNN 实现。Analytically computing the optimal representation in terms of a contrastive loss for a range of basic augmentations and any image dataset with stationary statistics shows that for certain augmentations the optimum can be attained by a CNN whose first layer filters are sinusoids.

MAGIC: Transition-Aware Generation of Navigable Multi-Scene Game Worlds with Large Language Models
MAGIC:基于 LLM 的转换感知可导航多场景游戏世界生成
arXiv:2607.11594 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

MAGIC 是一个四阶段 pipeline,能将单一自然语言提示转化为可运行的多场景游戏项目,相比 LLM 基线与 Holodeck 可恢复更多真实 portal,并生成显著更可导航的布局。MAGIC is a four-stage pipeline that turns a single natural-language prompt into a runnable multi-scene game project that recovers more ground-truth portals and yields markedly more navigable layouts than an LLM baseline and Holodeck.

Hallo4D: Multi-Modal Hallucination Mitigation for Consistent Spatio-Temporal Generation
Hallo4D: Multi-Modal Hallucination Mitigation for Consistent Spatio-Temporal Generation
arXiv:2607.12752 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

Hallo4D 提出"生成-检测-修正"范式,利用大型多模态语言模型(LMMs)从多视角与多帧渲染中识别并归纳时空不一致性,为一致性感知的内容生成提供了一种可扩展且可泛化的方案。Hallo4D introduces a generation-detection-correction paradigm that leverages large multimodal language models (LMMs) to identify and summarize spatial and temporal inconsistencies from multi-view and multi-frame renderings, providing a scalable and generalizable solution for consistency-aware content generation.

Boogu-Image-0.1: Boosting Open-Source Unified Multimodal Understanding and Generation
Boogu-Image-0.1: Boosting Open-Source Unified Multimodal Understanding and Generation
arXiv:2607.13125 多模态 方法 OA · 绿色 被引 1 · S2

研究表明,通过更强的多模态编码器、Agentic prompt 改写及相关技术来增强 Boogu-Image 系统的理解能力,并结合数据质量、训练流程和 Agentic 推理时扩展的改进,即使在计算预算极为受限的条件下,也能显著提升生成与编辑性能。It is demonstrated that strengthening the understanding capability of the Boogu-Image system, through a stronger multimodal encoder, agentic prompt rewriting, and related techniques, together with improvements in data quality, training pipelines, and agentic inference-time scaling, can substantially enhance generation and editing performance even under highly constrained compute budgets.

Registers Matter for Pixel-Space Diffusion Transformers
Registers 对像素空间 Diffusion Transformer 至关重要
arXiv:2605.16147 多模态 方法 OA · 绿色 被引 2 · S2

本研究表明 DiT 与 ViT 在一个关键方面存在差异:DiT 不会出现 patch-token 异常值,但仍能受益于 registers;并且 registers 在像素空间 DiT 中比在潜空间 DiT 中效果更显著。This work shows that DiTs differ from ViTs in a key respect: they do not exhibit patch-token outliers but still benefit from registers, and finds that registers are more effective in pixel-space DiTs than in latent-space DiTs.

AffectFlow-DINO: Uncertainty-Aware Multi-Task Affect Estimation via Conditional Rectified Flow
AffectFlow-DINO:基于条件 Rectified Flow 的不确定性感知多任务情感估计
arXiv:2607.13250 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

面向第 11 届 ABAW 挑战赛的多任务学习系统,在标准确定性架构基础上扩展条件 Rectified Flow 头,建模真实场景下面部行为固有的模糊性,借助蒙特卡洛采样实现不确定性感知的一对多预测。A multi-task learning system for the 11th ABAW challenge that extends a standard deterministic architecture with a conditional rectified-flow head to model the inherent ambiguity of in-the-wild facial behavior, enabling uncertainty-aware one-to-many predictions through Monte Carlo sampling.

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding
VideoChat3:面向高效通用视频理解的完全开放视频 MLLM
arXiv:2607.14935 多模态 应用落地 OA · 绿色 被引 4 · S2

本文提出 VideoChat3,一个完全开放、高效且以视频为中心的通用 MLLM,仅以 4B 参数与更高效率,超越参数量相当或更大的已有开源模型。This work introduces VideoChat3, a fully open, efficient, and generalist video-centric MLLM, which surpasses prior open-source models with equal or larger parameter counts with only 4B parameters and higher efficiency.

KeyFrame-Compass: Towards Comprehensive Evaluation of Keyframe-Conditioned Video Generation
KeyFrame-Compass:迈向关键帧条件视频生成的综合评估
arXiv:2607.14202 多模态 评测集 OA · 绿色 被引 1 · S2

本文将关键帧执行拆解为存在性、保真度、时序顺序、定位、持续性与唯一性六个互补指标,并通过结合专用感知模型的、基于证据的 MLLM 判断来评估整体视频质量。This work decomposes keyframe execution into six complementary metrics covering presence, fidelity, temporal ordering, localization, persistence, and uniqueness, while assessing overall video quality through evidence-grounded MLLM judgments augmented with specialized perception models.

MultiRef-Compass: Towards Comprehensive Evaluation of Multi-Reference-to-Audio-Video Generation
MultiRef-Compass:迈向多参考音视频生成任务的综合评估
arXiv:2607.14189 多模态 评测集 OA · 绿色 被引 1 · S2

本文提出 MultiRef-Compass,一个面向 MR2AV 生成的统一基准,将自动指标与引入复判增强的 MLLM-as-a-Judge 框架相结合,实现对感知保真度与参考条件合成能力的可扩展、可审计评估。MultiRef-Compass is introduced, a unified benchmark for MR2AV generation that integrates automatic metrics with a rejudging-enhanced MLLM-as-a-Judge framework, enabling scalable and auditable evaluation of both perceptual fidelity and reference-conditioned composition.

From Pixels to States: Rethinking Interactive World Models as Game Engines
从像素到状态:把交互式世界模型重新思考为游戏引擎
arXiv:2607.14076 多模态 观点 OA · 绿色 被引 2 · S2

本文从玩家动作控制、游戏状态动态、状态-观测持久性与实时交互生成四个维度审视交互式游戏世界建模,并针对《Black Myth: Wukong》提出可扩展的数据引擎,采集超过 90 小时的游戏画面作为状态感知型游戏世界建模的资源。This paper examines interactive game world modeling along four dimensions: player action control, game state dynamics, state-observation persistence, and real-time interactive generation, and presents a scalable data engine for Black Myth: Wukong that collects over 90 hours of gameplay as a resource for state-aware game world modeling.

VIABench: A Comprehensive Video Benchmark Collected from Blind Individuals for Visual Impairment Assistance
VIABench:面向视障辅助任务的盲人视频综合基准
arXiv:2607.14660 多模态 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

提出 VIABench,一个专为评估 MLLM 在视障辅助(VIA)场景中表现而设计的综合视频基准,采用视障人士(VIIs)自行录制或分享的第一人称视频,并提出一套严格的评测流水线,同时支持在线(实时)与离线设置。VIABench is introduced, a comprehensive video benchmark specifically designed to evaluate MLLMs in Visually Impaired Assistance scenarios using first-person videos recorded or shared by VIIs themselves, and proposes a rigorous benchmarking pipeline that supports both online (real-time) and offline settings.

Hierarchical Denoising For Multi-Step Visual Reasoning
面向多步视觉推理的分层去噪方法
arXiv:2607.15278 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 HDR (Hierarchical Denoising for Visual Reasoning),一个将层级潜变量集成到因果视频生成中以进行多步推理的统一框架,并引入一个含分布外情况的层级化多步视频推理基准。This work proposes HDR (Hierarchical Denoising for Visual Reasoning), a unified framework that integrates hierarchical latents into causal video generation for multi-step reasoning and introduces a level-stratified multi-step video reasoning benchmark with out-of-distribution cases.

Flamingo: a Visual Language Model for Few-Shot Learning
Flamingo: a Visual Language Model for Few-Shot Learning
arXiv:2204.14198 多模态 方法 OA · 绿色 被引 6453 · S2

提出 Flamingo,一个 VLM 系列,能够桥接强大的纯视觉与纯语言预训练模型,处理任意交错排列的图文序列,并无缝接收图像或视频作为输入。This work introduces Flamingo, a family of Visual Language Models (VLM) with this ability to bridge powerful pretrained vision-only and language-only models, handle sequences of arbitrarily interleaved visual and textual data, and seamlessly ingest images or videos as inputs.

A Comprehensive Survey of AI-Generated Content (AIGC): A History of Generative AI from GAN to ChatGPT
A Comprehensive Survey of AI-Generated Content (AIGC): A History of Generative AI from GAN to ChatGPT
arXiv:2303.04226 多模态 综述 OA · 绿色 被引 838 · S2

该综述全面回顾了生成模型的历史与基本组件,以及 AIGC 在单模态交互与多模态交互方向的最新进展,并介绍了文本与图像生成任务及相关模型。This survey provides a comprehensive review on the history of generative models, and basic components, recent advances in AIGC from unimmodal interaction and multimodal interaction, and introduces the generation tasks and relative models of text and image.

Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
arXiv:2403.05530 多模态 方法 OA · 绿色 被引 3916 · S2

Gemini 1.5 在跨模态长上下文检索任务上取得近乎完美的召回率,在长文档 QA、长视频 QA 与长上下文 ASR 上刷新 SOTA,并在广泛基准上达到或超越 Gemini 1.0 Ultra 的 SOTA 表现。Gemini 1.5 models achieve near-perfect recall on long-context retrieval tasks across modalities, improve the state-of-the-art in long-document QA, long-video QA and long-context ASR, and match or surpass Gemini 1.0 Ultra's state-of-the-art performance across a broad set of benchmarks.

A Review on Generative Adversarial Networks: Algorithms, Theory, and Applications
A Review on Generative Adversarial Networks: Algorithms, Theory, and Applications
arXiv:2001.06937 多模态 综述 OA · 绿色 被引 1154 · S2

详细介绍大多数 GAN 算法的动机、数学表示与结构,并对它们的共性与差异进行比较。The motivations, mathematical representations, and structures of most GAN algorithms are introduced in detail, and they are compared to compare their commonalities and differences.

On Locality and Length Generalization in Visual Reasoning
关于视觉推理中的局部性与长度泛化
arXiv:2607.09061 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

实验表明,与语言模型类似,视觉模型也会学习利用全局捷径,从而无法在任务长度或复杂度上泛化,但本文证明基于严格局部感知的循环视觉策略可以缓解这些失败,从而使模型在这些任务上具备泛化能力。The experiments reveal that, similar to language models, vision models can learn to exploit global shortcuts and thereby fail to generalize over task length or complexity, but it is shown that recurrent vision policies based on strictly local perception can mitigate these failures, thereby allowing models to generalize on these tasks.

AsySplat: Efficient Asymmetric 3D Gaussian Splatting for Long-Sequence Scene Modeling
AsySplat:用于长序列场景建模的高效非对称 3D Gaussian Splatting
arXiv:2607.10995 多模态 方法 被引 0 · S2

该模型在效果上匹配基于优化的方法,同时实现近 800 倍加速,并以明显更少的参数量和更低的训练/推理开销超越 SOTA 可泛化模型的零样本性能,整体效率显著提升。This model matches optimization-based methods while delivering nearly 800x speedup, and surpasses the zero-shot performance of state-of-the-art generalizable models with markedly fewer parameters and reduced training/inference overhead, achieving an overall efficiency improvement.

Token Time Continuous Diffusion for Language Modeling
用于语言建模的 Token 时间连续扩散
arXiv:2607.14106 多模态 方法 被引 0 · S2

本文提出 token 时间连续扩散 (TTCD),一种新的扩散语言模型,运行于连续空间,将高斯噪声确定性地映射到最终的 token canvas 而无需额外采样,并引入每个 token 时间的新概念。This paper introduces token time continuous diffusion (TTCD), a new diffusion language model which operates in continuous space, deterministically mapping Gaussian noise to a final token canvas with no further sampling, and incorporates a new notion of per-token times.

S1-Omni: A Unified Multimodal Reasoning Model for Scientific Understanding, Prediction, and Generation
S1-Omni:用于科学理解、预测与生成的统一多模态推理模型
arXiv:2607.15686 多模态 方法 被引 0 · S2

S1-Omni 基于 S1-Omni-Corpus 训练,覆盖 200 个科学任务并包含数百万推理样本,在 60 余个科学基准上进行了评估,为统一的科学建模提供了一条可行路径。S1-Omni is trained on S1-Omni-Corpus, which covers 200 scientific tasks and contains millions of reasoning samples, and is evaluated on over 60 scientific benchmarks, providing a practical path toward unified scientific modeling.

On-Policy Delta Distillation
在策略差值蒸馏
arXiv:2607.15161 多模态 方法 被引 2 · S2

表明差值信号显著提升在策略蒸馏效果,提出新方法称为在策略差值蒸馏(OPD),使推理LLM仅需短暂的后训练即可获得强性能。It is shown that the delta signal substantially improves on-policy distillation and the new distillation method is referred to as On-Policy Delta Distillation (OPD), enabling reasoning LLMs to achieve strong performance with only a short post-training period.

VideoRAE: Taming Video Foundation Models for Generative Modeling via Representation Autoencoders
VideoRAE:通过表征自编码器驯服视频基础模型用于生成建模
arXiv:2607.14088 多模态 方法 被引 0 · S2

VideoRAE是一种表征自编码器,利用冻结视频基础编码器的多尺度分层特征,并通过轻量级1D自注意力投影器进行压缩,验证了冻结VFM表征可作为通用且利于生成的视频潜变量。VideoRAE is a representation autoencoder that leverages multi-scale hierarchical features from a frozen video foundation encoder and compresses them with a lightweight 1D self-attention projector to validate frozen VFM representations as versatile and generation-friendly video latents.