Papers · organized/paper_cards

论文

258 张论文卡片 · 多模态

开放获取 全部 绿色 · 724
Invisible Shortcuts: Why Vision Encoders Know Your Camera
Invisible Shortcuts:视觉 Encoder 为何知道你的相机
arXiv:2608.05424 多模态 方法 被引 0 · S2

识别出像素级嵌入的不可见元数据痕迹,表明大规模语义监督(无论是类别标签还是亿级规模的 caption)在预训练中会自然引发元数据-语义相关性,导致模型将低层信号转化为预测特征。Invisible metadata traces embedded at the pixel level are identified, suggesting that large-scale semantic supervision, whether through categorical labels or billion-scale captions, naturally induces metadata-semantics correlations during pretraining, leading models to convert low-level signals into predictive features.

MASS: Multiplayer World Models with Authoritative Shared State
MASS:基于权威共享状态的多智能体世界模型
arXiv:2608.06257 多模态 方法 被引 0 · S2

结果表明,显式且权威的状态建模为可扩展、一致的多智能体世界仿真提供了可行基础。The results show that explicit, authoritative state modeling provides a practical foundation for scalable and consistent multi-agent world simulation.

Transformer in Transformer
Transformer in Transformer
arXiv:2103.00112 多模态 方法 OA · 绿色 被引 2262 · S2

本工作指出,这些局部 patch 内部的注意力同样是构建高性能视觉 Transformer 的关键,并探索了一种新架构,即 Transformer iN Transformer (TNT)。It is pointed out that the attention inside these local patches are also essential for building visual transformers with high performance and a new architecture, namely, Transformer iN Transformer (TNT), is explored.

KVAE: Family of Tokenizers for Multimodal Generative Models
KVAE:面向多模态生成模型的 tokenizer 家族
arXiv:2608.05798 多模态 观点 被引 0 · S2

在客观与主观指标上的重建和生成结果匹配或超越前沿开源 tokenizer,包括 Wan-2.2、HunyuanVideo-1.5、FLUX.2、MovieGen、StableAudio 和 MMAudio 的 VAE。It is demonstrated that reconstruction and generation results on objective and subjective metrics matches or surpasses frontier opensource tokenizers, such as VAEs from Wan-2.2, HunyuanVideo-1.5, FLUX.2, MovieGen, StableAudio and MMAudio.

Weights or Skills? A Survey of Robot-Learning Techniques: from Action-Predicting Weights to Robots that Write their Own Skills
Weights 还是 Skills?机器人学习方法综述:从预测动作的权重到自主编写技能的机器人
arXiv:2608.01851 多模态 综述 被引 0 · S2

该综述围绕"权重"与"技能"这一轴线组织领域,梳理了互补的"技能"一极——从无监督强化学习的技能发现,到大语言模型的技能库——并指出"skill"一词至少存在五种不同含义。This survey organises the field around that axis of weights versus skills, and maps the complementary"skills"pole, from unsupervised reinforcement-learning skill discovery to large-language-model skill libraries, and shows that the word "skill" is used in at least five distinct senses.

Representation Learning with Contrastive Predictive Coding
基于对比预测编码的表征学习
arXiv:1807.03748 多模态 方法 OA · 绿色 被引 14265 · S2

本文提出了一种通用的无监督学习方法——对比预测编码(Contrastive Predictive Coding),用于从高维数据中提取有用的表征,并在语音、图像、文本和 3D 环境中的强化学习四个不同领域取得了出色的性能。This work proposes a universal unsupervised learning approach to extract useful representations from high-dimensional data, which it calls Contrastive Predictive Coding, and demonstrates that the approach is able to learn useful representations achieving strong performance on four distinct domains: speech, images, text and reinforcement learning in 3D environments.

SimWAM: A Simple World Action Model for End-to-End Autonomous Driving
SimWAM:用于端到端自动驾驶的简单 World Action Model
arXiv:2608.07468 多模态 方法 被引 0 · S2

SimWAM 是一种简洁而有效的 WAM,利用未来视频预测作为训练期监督信号,通过 joint flow matching 联合训练一个预训练视频专家与一个轻量级动作专家,并采用强化学习在轨迹模仿之上优化组合式驾驶奖励。SimWAM, a simple yet effective WAM that leverages future-video prediction as a training-time supervision signal, co-trains a pretrained video expert and a lightweight action expert with joint flow matching and applies reinforcement learning to optimize a compositional driving reward beyond trajectory imitation.

Round-Trip Consistency: Bidirectional Diffusion Models Can Predict Their Own Rollout Errors
往返一致性:双向扩散模型可预测自身的 rollout 误差
arXiv:2608.00675 多模态 应用落地 被引 1 · S2

往返一致性将可逆性转化为生成式模型一种实用的可信信号;双向训练带来负成本,在两个方向上均优于单向专家模型;其中反向还可作为快速的逆问题求解器。Round-trip consistency turns reversibility into a practical trust signal for generative models, and Bidirectional training comes at negative cost, beating direction specialists in both directions, and the backward direction doubles as a fast inverse solver.

Diffusion-Convolutional Neural Networks
扩散卷积神经网络
arXiv:1511.02136 多模态 方法 OA · 绿色 被引 1379 · S2

通过引入扩散卷积运算,本文展示了如何从图结构数据中学习基于扩散的表示,并将其作为节点分类的有效基础。Through the introduction of a diffusion-convolution operation, it is shown how diffusion-based representations can be learned from graph-structured data and used as an effective basis for node classification.

Florence: A New Foundation Model for Computer Vision
Florence:面向计算机视觉的新基础模型
arXiv:2111.11432 多模态 方法 OA · 绿色 被引 1152 · S2

本文提出新的计算机视觉基础模型 Florence,通过融入来自 Web 规模图文数据的通用视觉-语言表示,将表征范围从粗粒度(场景)扩展到细粒度、从静态(图像)扩展到动态(视频)、从 RGB 扩展到多种模态(描述、深度等)。This work introduces a new computer vision foundation model, Florence, to expand the representations from coarse (scene) to fine, from static (images) to dynamic (videos), and from RGB to multiple modalities (caption, depth), by incorporating universal visual-language representations from Web-scale image-text data.

Co-occurrence Feature Learning for Skeleton based Action Recognition using Regularized Deep LSTM Networks
使用正则化深度 LSTM 网络进行基于骨骼动作识别的共现特征学习
arXiv:1603.07772 多模态 应用落地 OA · 绿色 被引 930 · S2

本文在每个时间步以骨骼作为输入,引入一种新的正则化方案来学习骨骼关节的共现特征,并提出一种同时作用于 LSTM 神经元门、单元和输出响应的新型 dropout 算法。This work takes the skeleton as the input at each time slot and introduces a novel regularization scheme to learn the co-occurrence features of skeleton joints, and proposes a new dropout algorithm which simultaneously operates on the gates, cells, and output responses of the LSTM neurons.

Multi-Agent Forensic Reasoning for Generalizable Deepfake Video Detection
多智能体取证推理用于可泛化的深度伪造视频检测
arXiv:2608.06865 多模态 评测集 被引 0 · S2

提出 FaceVid-Forensics-100K,一个大规模深伪视频数据集,包含 100,000 个视频,涵盖 33 种合成方法,覆盖换脸、表情重演与全脸合成;同时提出一个多智能体取证推理框架,由四个领域专家 Agent 分别从四个角度独立分析伪造线索。FaceVid-Forensics-100K is introduced, a large-scale deepfake video dataset comprising 100,000 videos and spanning 33 synthesis methods across face swapping, face reenactment, and entire-face synthesis, and a multi-agent forensic reasoning framework that employs four specialized domain-expert agents to independently analyze forgery cues from four perspectives.

Enfold: Folding World Model Imagination into Predictive Representations for Ultra-Efficient Embodied Control
Enfold:将世界模型想象折叠进预测表征以实现超高效具身控制
arXiv:2607.26657 多模态 方法 被引 0 · S2

本文提出 Enfold,将构建未来的计算转移到由当前视觉上下文和语言指令预测出的表征中,并把世界生成器重塑为预测控制表征的来源,前提是其内部结构可被折叠(enfold)到当下。This work presents Enfold, which transfers this computation that constructs a future into a representation predicted from the current visual context and language instruction, and recast a world generator as a source of predictive control representations if its internal structure can be enfolded into the present.

CLIP-CC-Bench: Evaluating Paragraph-Level Video Descriptions in Video-Language Models
CLIP-CC-Bench:评估视频-语言模型中的段落级视频描述
arXiv:2608.04302 多模态 评测集 被引 0 · S2

CLIP-CC-Bench 为长视频描述提供了一个实用的评估框架,填补了现有短片段和仅 QA 基准的空白,并通过评分者间一致性(inter-judge agreement)与 bootstrap 排序稳定性量化该协议的内可靠性。CLIP-CC-Bench provides a practical evaluation framework for long-form video description, filling a gap left by existing short-clip and QA-only benchmarks and quantifying the protocol's internal reliability through inter-judge agreement and bootstrap ranking stability.

DuplexGen: Adaptive Synthesis of Human-AI Turn-Taking Dialogues
DuplexGen:自适应合成人-AI 轮替对话
arXiv:2607.26178 多模态 方法 被引 0 · S2

本文提出 DuplexGen,一个通过少量 slot 级人类偏好标注对 LLM 预测进行校准从而生成具有场景自适应轮换特性的对话框架;结果表明,使轮换合成具备场景特异性的关键是人类校准,而非单纯的语料规模或提示设计。DuplexGen is introduced, a framework for generating dialogues with scenario-adaptive turn-taking by calibrating LLM predictions against a small set of slot-level human preference annotations, and results show that human calibration, not corpus scale or prompt design alone, is what allows turn-taking synthesis to be scenario-specific.

RynnValue: Scaling Robotic Value Foundation Models with Temporal Distance
RynnValue:通过时间距离扩展机器人价值基础模型
arXiv:2608.09853 多模态 方法 被引 1 · S2

RynnValue 是一个用于机器人操控的开源价值基础模型,用时间距离(即从观测到语言指定目标的有向 cost-to-go)替代任务内部锚点,将时间距离确立为通用机器人策略的可扩展监督目标和实用奖励接口。RynnValue, an open-source value foundation model for robotic manipulation that replaces task-internal anchors with temporal distance, the directed cost-to-go from an observation to the language-specified goal, establishes temporal distance as a scalable supervision target and practical reward interface for generalist robot policies.

Scaling Inherently Interpretable Language Models
规模化天生可解释的语言模型
arXiv:2608.07594 多模态 方法 被引 2 · S2

Steerling-8B 在与训练计算量多 2–16 倍的开源同侪模型对比中仍保持竞争力,表明存在一种不同的可扩展范式:可解释性可以被设计进训练过程中,并随规模放大而提升。Steerling-8B remains competitive with open peer models trained on substantially 2-16x more compute, suggesting a different scaling paradigm: interpretability can be designed into training, and it improves with scale.

What to Edit Next: Visually Aligned Image-Editing Follow-Up Suggestions in Conversational Systems
下一步编辑什么:对话系统中视觉对齐的图像编辑后续建议
arXiv:2608.07565 多模态 方法 被引 0 · S2

一个三阶段框架,用于减少建议编辑与当前图像之间的视觉不一致性,并将推荐 CTR、图像带走率以及用户平均对话轮次分别显著提升 39.90%(所有 p<0.05)。A three-stage framework to reduce visual inconsistencies between suggested edits and the current image, which significantly improves recommendation CTR, image take-away rate, and average conversation turns per user by 39.90% (all p<0.05).

Ego-OSCAR: Egocentric Open source Stereo CAptuRe System
Ego-OSCAR:自我中心开源立体捕获系统
arXiv:2608.08285 多模态 方法 被引 0 · S2

本文提出 Ego-OSCAR,一种用于野外自我中心数据采集的开源硬件、低成本、头戴式立体惯性采集设备,旨在成为众包自我中心采集中最廉价且可辩护的载体,降低任何团队大规模采集自我中心数据的启动门槛。Ego-OSCAR is presented, an open-hardware, low-cost, head-mounted stereo-inertial capture device for egocentric data collection in the wild that aims to be the cheapest defensible substrate for crowdsourced egocentric capture, and to lower the activation energy for any team that wants to collect egocentric data at scale.

Vision-Language Grounding as Bidirectional Concept Correspondence
视觉-语言接地作为双向概念对应
arXiv:2608.07886 多模态 方法 被引 0 · S2

该形式化方法将短语定位、指代表达定位以及开放词汇检测等常见 grounding 任务统一起来,将文本分割、图像分割以及跨模态对齐视为单一的对应预测问题。This formulation unifies common grounding tasks, including phrase grounding, referring expression grounding, and open-vocabulary detection, by treating text segmentation, image segmentation, and cross-modal alignment as a single correspondence prediction problem.

Domain Adaptation for Visual Applications: A Comprehensive Survey
视觉应用中的域适应:综合综述
arXiv:1702.05374 多模态 综述 OA · 绿色 被引 554 · S2

综述领域自适应与迁移学习,重点关注视觉应用及超越图像分类的方法,如目标检测、图像分割、视频分析或视觉属性学习。An overview of domain adaptation and transfer learning with a specific view on visual applications and the methods that go beyond image categorization, such as object detection or image segmentation, video analyses or learning visual attributes are overviewed.

Beyond Starry Night: Shortcut-Aware Control-State Planning for Artist-Grounded Text to Image Generation
Beyond Starry Night: Shortcut-Aware Control-State Planning for Artist-Grounded Text to Image Generation
arXiv:2608.06751 多模态 方法 被引 0 · S2

Atelier 是一种面向艺术家风格图像生成的捷径感知控制状态规划框架,可提升艺术家级风格保真度,更忠实地保持源结构,并相较提示工程、检索增强与通用 agent 基线大幅减少捷径替换。Atelier, a shortcut-aware control-state planning framework for artist-grounded image generation, improves artist-level style fidelity, preserves source structure more faithfully, and substantially reduces shortcut substitution compared with prompt-engineered, retrieval-augmented, and general-purpose agent baselines.

Hand Visibility Detector: Per-Keypoint Visibility Estimation for Hands
手部可见性检测器:基于关键点的手部可见性估计
arXiv:2608.11574 多模态 方法 被引 0 · S2

本文表明,利用在大规模数据上预训练的 HPE 模型先验知识作为骨干网络可在该任务上取得高性能,并验证了手部可见性检测器在通过 2D 关键点多视角三角化进行 3D 手部姿态标注的下游任务中的有效性。It is shown that leveraging the prior knowledge of HPE models pretrained on large-scale data as a backbone yields high performance in this task, and the utility of Hand Visibility Detector on a downstream task of 3D hand pose annotation via multi-view triangulation of 2D keypoints.

AtlasVLA: Persistent World-Ego State Modeling for Vision-Language-Action Models
AtlasVLA:面向视觉—语言—动作模型的持久世界—自我状态建模
arXiv:2608.06729 多模态 方法 被引 0 · S2

AtlasVLA 是一种新框架,通过持久化的世界-自我状态从直接反应式操作转向主动推理,显著优于多视角基线,在 LIBERO-Long 上取得 9.4% 的绝对成功率提升,在真实世界长周期任务中取得 17.5% 的提升。AtlasVLA is a novel framework that transitions from direct reactive manipulation to proactive reasoning through a persistent world-ego state and decisively outperforms multi-view baselines, yielding absolute success rate improvements of 9.4% on LIBERO-Long and 17.5% in real-world long-horizon tasks.

Simplex Relaxation for Discrete Diffusion
离散扩散的单纯形松弛
arXiv:2608.10615 多模态 方法 被引 0 · S2

本文提出 Simplax,一种精确的 Dirichlet-类别增强方法,将每个被破坏的类别状态与一个辅助的 simplex 值变量耦合,同时保留原始均匀扩散过程作为其类别边缘分布。Simplax is introduced, an exact Dirichlet--categorical augmentation that couples each corrupted categorical state with an auxiliary simplex-valued variable while preserving the original uniform diffusion process as its categorical marginal.

Gaze Target Estimation Anywhere with Concepts
Gaze Target Estimation Anywhere with Concepts
arXiv:2608.11367 多模态 方法 被引 2 · S2

本文提出提示式注视目标估计(PGE)任务,一种用于注视分析的端到端、概念驱动新范式,并推出首个为 PGE 设计的模型 GazeAnywhere,它使用基于 Transformer 的检测器融合冻结编码器特征,同时解决主体定位、画内/画外存在性以及注视目标热图估计。The Promptable Gaze Target Estimation (PGE) task is introduced, a new end-to-end, concept-driven paradigm for gaze analysis and GazeAnywhere, the first model designed for PGE, uses a transformer-based detector to fuse features from frozen encoders and simultaneously solves subject localization, in/out-of-frame presence, and gaze target heatmap estimation.

DreamX-Phi 1.0: Action-Conditioned Video World Model for Robotic Manipulation
DreamX-Phi 1.0:面向机器人操作的动作条件视频世界模型
arXiv:2608.13489 多模态 方法 被引 0 · S2

本文提出一种用于机器人操作的动作条件视频世界模型,给定观测帧、语言指令以及由末端执行器位姿和夹爪状态组成的预定动作序列,预测对应的未来观测结果。An action-conditioned video world model for robotic manipulation that, given an observed frame, a language instruction, and a prescribed action sequence comprising end-effector poses and gripper states, predicts the resulting future observations is presented.

UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos
UniSwap:面向说话视频的流式音视频身份替换
arXiv:2608.11752 多模态 方法 被引 0 · S2

本文提出 UniSwap,这是首个用于说话视频中流式联合音视频身份替换的框架,并引入 swap-and-reconstruct 流程,从真实片段中移除视觉和声音身份,同时使用原始片段作为重建目标。This work presents UniSwap, the first framework for streaming joint audio-visual identity replacement in talking videos, and introduces a swap-and-reconstruct pipeline that removes visual and vocal identity from real clips and uses the original clips as reconstruction targets.

LiveAnimate: Stable Long-Form Streaming Human Animation in Real-Time
LiveAnimate:实时稳定的长视频流式人体动画
arXiv:2608.11745 多模态 应用落地 被引 0 · S2

本文提出 LiveAnimate,据作者所知是首个将实时流式生成与十亿参数规模下的稳定长视频生成相结合的系统,基于 140 亿参数的视频 Diffusion Transformer(DiT)。This work presents LiveAnimate, to their knowledge the first animation system to combine real-time streaming with stable long-form generation at billion scale, built on a 14B-parameter video Diffusion Transformer (DiT).

H2R-Bench: Benchmarking Human-to-Robot Manipulation Video Generation in World Models
H2R-Bench:世界模型中人到机器人操作视频生成的基准评测
arXiv:2608.13049 多模态 评测集 被引 1 · S2

H2R-Bench 提供了一个系统性诊断框架,用于评估视频世界模型能否跨越 human-to-robot 具身差距,并将人类操作观测转化为以机器人为中心的训练资源。H2R-Bench provides a systematic diagnostic framework for evaluating whether video world models can bridge the human-to-robot embodiment gap and convert human manipulation observations into robot-centric training resources.

An AI4AI Framework for Visual Token Pruning
面向视觉 token 剪枝的 AI4AI 框架
arXiv:2608.07193 多模态 方法 被引 0 · S2

本文认为关键在于设计合适的 search-state 表示,将 LLM 的内部知识与 visual-token 剪枝的结构要求和约束相连接,并提出 AutoPrune,一种用于 LLM 驱动的 visual-token 剪枝策略设计的免训练框架。This paper argues that the key lies in designing an appropriate search-state representation that connects the internal knowledge of LLMs with the structural requirements and constraints of visual-token pruning, and proposes AutoPrune, a training-free framework for LLM-driven visual-token pruning policy design.

Context-Matched Distillation: Teacher Causality for Autoregressive Video Distillation
上下文匹配蒸馏:用于自回归视频蒸馏的教师因果性
arXiv:2608.13391 多模态 方法 被引 0 · S2

提出 Context-Matched Distillation(CMD),一种因果 DMD 框架,将 teacher 监督信号与每个 target 生成时可用信息对齐,并可自然扩展到逐帧和逐 chunk 生成、长视频蒸馏以及相机条件蒸馏。Context-Matched Distillation (CMD) is introduced, a causal DMD framework that aligns teacher supervision with the information available when each target is generated, and naturally extends to frame-wise and chunk-wise generation, long video distillation, and camera-conditioned distillation.

Robust Speech Recognition via Large-Scale Weak Supervision
Robust Speech Recognition via Large-Scale Weak Supervision
arXiv:2212.04356 多模态 方法 OA · 绿色 被引 8184 · S2

当将监督规模扩展到 680,000 小时的多语言、多任务数据时,所得到的模型在标准 benchmark 上泛化良好,在 zero-shot transfer 设置下常可与此前全监督方法的结果相当,且无需任何微调。When scaled to 680,000 hours of multilingual and multitask supervision, the resulting models generalize well to standard benchmarks and are often competitive with prior fully supervised results but in a zero-shot transfer setting without the need for any fine-tuning.

RibAssist 3D: Biplanar Rib-Fracture Detection, Addressing, and Selective 3D Localization from CT-Derived Projections
RibAssist 3D:基于 CT 投影的双平面肋骨骨折检测、配对与选择性 3D 定位
arXiv:2608.06914 多模态 方法 被引 0 · S2

建立一个可复现的选择性 3D 定位框架,并指出跨视角对应(cross-view correspondence)是主要操作瓶颈。The study establishes a reproducible framework for selective 3D localization and identifies cross-view correspondence as the dominant operational bottleneck.

Multimodal Model Diffing for Feature Discovery and Control
多模态模型 Diffing:用于特征发现与控制
arXiv:2608.09928 多模态 方法 被引 0 · S2

MMDiff 是一种多模态 model-diffing 框架,训练多模态 SAE 并将其转化为特征级接口,用于发现和控制多模态行为;研究表明多模态 SAE 不仅可作为可解释性工具,还可作为审计、引导和控制 MLLM 行为的机制,以实现更安全、更具能力的生成。MMDiff, a multimodal model-diffing framework that trains multimodal SAEs and turns them into feature-level interfaces for discovering and controlling multimodal behavior, shows that multimodal SAEs can serve not only as interpretability tools, but as mechanisms for auditing, steering, and controlling MLLMs behavior toward safer and more capable generations.

A Pathway to General-Purpose Scientific AI: Multimodal Comprehension of Scientific Images
迈向通用科学 AI 的路径:科学图像的多模态理解
arXiv:2608.14075 多模态 评测集 被引 1 · S2

展望 ALD/E-ImageMiner benchmark 如何指导未来科学图像挑战赛,以及 Bloom-informed 问题设计如何支持更深入的科学理解。A forward-looking perspective is presented on how the ALD/E-ImageMiner benchmark can guide future scientific-image challenges and how Bloom-informed question design can support deeper scientific understanding.