Papers · organized/paper_cards

论文

258 张论文卡片 · 多模态

开放获取 全部 绿色 · 724
PointNet: Deep Learning on Point Sets for 3D Classification and Segmentation
PointNet: Deep Learning on Point Sets for 3D Classification and Segmentation
arXiv:1612.00593 多模态 方法 OA · 绿色 被引 18312 · S2

本文设计了一种直接处理点云的新型神经网络,较好地尊重了输入点的置换不变性,并为从物体分类、部件分割到场景语义解析等应用提供了统一架构。This paper designs a novel type of neural network that directly consumes point clouds, which well respects the permutation invariance of points in the input and provides a unified architecture for applications ranging from object classification, part segmentation, to scene semantic parsing.

Gemini: A Family of Highly Capable Multimodal Models
Gemini: A Family of Highly Capable Multimodal Models
arXiv:2312.11805 多模态 方法 OA · 绿色 被引 829 · OpenAlex
Synthetic Data and Artificial Neural Networks for Natural Scene Text Recognition
Synthetic Data and Artificial Neural Networks for Natural Scene Text Recognition
arXiv:1406.2227 多模态 方法 OA · 绿色 被引 1000 · S2

本文提出了一个自然场景文本识别框架,无需任何人工标注数据,并以整体方式对整幅图像进行单词识别,区别于过去基于字符的识别系统。This work presents a framework for the recognition of natural scene text that does not require any human-labelled data, and performs word recognition on the whole image holistically, departing from the character based recognition systems of the past.

ENTRAP-VL: A Taxonomic Probe for Dual Contextual Entrainment in Vision-Language Models
ENTRAP-VL: A Taxonomic Probe for Dual Contextual Entrainment in Vision-Language Models
arXiv:2607.20092 多模态 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

提出 ENTRAP-VL(ENTRainment Assessment Probe for Vision and Language),一个由人工策展的 1,500 条数据的数据集,涵盖八个类别,按一个跨双轴的分类体系组织,并划分为文本诱发流和视觉诱发流。ENTRAP-VL (ENTRainment Assessment Probe for Vision and Language), a manually curated dataset of 1,500 items across eight categories, organized by a taxonomy that spans two axes and split into a textual-entrainment stream and a visual-entrainment stream, is introduced.

Recurrent Sinusoidal INRs for Efficient High-Fidelity Representation
循环正弦 INR 用于高效高保真表示
arXiv:2607.21485 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

研究揭示,正弦激活会诱发谐波线谱,为循环展开如何丰富隐式神经表示(INR)的有效频谱支撑提供了频谱层面的解释。It is revealed that sinusoidal activations induce a harmonic line spectrum, providing a spectral account of how recurrent unrolling enriches the effective spectral support in implicit neural representations (INRs).

ReferTrack: Referring Then Tracking for Embodied Visual Tracking
ReferTrack:先指代再跟踪的具身视觉跟踪
arXiv:2607.20061 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出 ReferTrack,一种"先指代后跟踪"的范式,仅使用单个前向摄像头完成 EVT grounding,在四足机器人和人形机器人上的真实部署验证了其稳健的 sim-to-real 迁移能力。ReferTrack is introduced, a referring-then-tracking paradigm that grounds EVT using a single forward-facing camera, and real-world deployments on legged and humanoid robots validate its robust sim-to-real transfer capabilities.

Robostral Navigate
Robostral Navigate
arXiv:2607.20785 多模态 应用落地 OA · 绿色 被引 0 · S2 + OpenAlex

提出 Robostral Navigate,一个围绕该可扩展性目标构建的 8B 视觉语言模型,仅消费单目 RGB 图像流——这是机器人平台中最普及的传感器——通过在当前相机画面中指向下一目标位置来预测航点。Robostral Navigate, an 8B vision-language model built around this scalability objective, is introduced, which consumes only a stream of monocular RGB images - the most ubiquitous sensor across robotic platforms and predicts waypoints by pointing to the next target location in the current camera view.

Color Pass-Through via Camera-Display Coupling
通过相机-显示器耦合实现色彩直通
arXiv:2607.12746 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 Color Pass-Through,一个端到端可学习的框架,直接在采集图像上运行,将相机和显示器视为耦合系统进行联合处理,而非单独校准。This work proposes Color Pass-Through, an end-to-end learned framework that operates directly on captured images, to treat the camera and display as a coupled system rather than calibrating them in isolation.

Self-Supervised Learning of Structured Dynamics from Videos
从视频中自监督学习结构化动态
arXiv:2607.21576 多模态 观点 被引 0 · S2

提出结构化动态模型(SDM),通过未来特征预测,显式地将时间变化的主导来源与残差动态分离开来,而非使用单一纠缠的隐变量或非结构化的、空间密集的转移 token 来表示视频变化。The Structured Dynamics Model (SDM) is proposed, which explicitly separates the dominant source of temporal change from residual dynamics through future-feature prediction, rather than representing video change with a single entangled latent or with unstructured, spatially dense transition tokens.

SANA-Video 2.0: Hybrid Linear Attention with Attention Residuals for Efficient Video Generation
SANA-Video 2.0:基于注意力残差的混合线性注意力高效视频生成
arXiv:2607.21553 多模态 方法 被引 0 · S2

本文提出 SANA-Video 2.0,一个混合视频扩散 Transformer,在统一架构下实例化 5B 和 14B 两种规模,以显著降低的计算成本恢复了 softmax 级别的表达能力,解锁了可扩展的长时长、高分辨率视频生成。This work introduces SANA-Video 2.0, a hybrid video diffusion transformer instantiated at 5B and 14B scales under a unified architecture that recovers softmax-level expressiveness at substantially reduced cost, unlocking scalable long, high resolution video generation.

Code of "Sirens' Whisper: Inaudible Near-Ultrasonic Jailbreaks of Speech-Driven LLMs"
《Sirens' Whisper:语音驱动 LLM 的不可听近超声越狱》代码
arXiv:2307.15043 多模态 方法 OA · 绿色 被引 3560 · S2

本文显著推进了针对已对齐语言模型的对抗攻击 SOTA,并提出了关于如何防止此类系统生成不良信息的重要问题。This work significantly advances the state-of-the-art in adversarial attacks against aligned language models, raising important questions about how such systems can be prevented from producing objectionable information.

Generative Adversarial Networks in Computer Vision: A Survey and Taxonomy
计算机视觉中的生成对抗网络:综述与分类
arXiv:1906.01529 多模态 综述 OA · 绿色 被引 277 · OpenAlex

深入回顾文献中 GAN 相关研究,并从两个视角阐述针对三大挑战所提出的架构变体与损失变体。An in-depth review of GAN-related research in the literature is provided, and an account of the architecture-variant and loss-variants, which have been proposed to handle these three challenges from two perspectives are provided.

Scaling Native Multimodal Pre-Training From Scratch
Scaling Native Multimodal Pre-Training From Scratch
arXiv:2607.22043 多模态 方法 被引 1 · S2

本实证研究通过建模数据组成对计算定律及分配指数的影响,推导出指定模型规模、token 数和数据混合精确配置的效率前沿,为可预测地扩展多模态基础模型奠定了必要基础。This empirical research establishes the essential groundwork for predictably scaling multimodal foundation models by modeling the influence of data composition on compute laws and allocation exponents and derive an efficiency frontier specifying precise configurations of model size, token count, and data mixture.

Closing the Loop: Training-Free Revisit Consistency for Autoregressive Generative Rendering
Closing the Loop:面向自回归生成式渲染的无训练 Revisit 一致性
arXiv:2607.21848 多模态 应用落地

近期条件视频生成模型已展现出将 3D 引擎渲染(如深度图与无纹理几何体)转化为照片级真实视频的潜力,可应用于游戏与沉浸式内容创作。此类应用要求长时程自回归生成,在持续合成新帧的同时维持持久的 3D 世界。自回归生成器以有界 KV cache 逐 chunk 合成视频,因此当相机再次访问已从上下文中驱逐的位置时,模型常会重新生成不一致的外观,尽管该Recent conditional video generation models have shown promising potentials to transform 3D engine renderings, such as depth maps and untextured geometry, into photorealistic videos for gaming and immersive content creation. These applications require long-horizon auto-regressive generation that continuously synthesizes new frames while preserving a persistent 3D world. Auto-regressive generators synthesize video chunk by chunk with a bounded KV cache, so when the camera revisits a location after its context has been evicted, the model often regenerates inconsistent appearance, even though the

VisCo: Leveraging Large Language Models as Intrinsic Encoders for Visual Token Compression
VisCo:利用大语言模型作为视觉 token 压缩的内在编码器
arXiv:2607.12756 多模态 方法 被引 0 · S2

VisCo 是一个训练高效的自压缩框架,复用预训练 VLM 本身作为内在压缩器,使用少量 memory token 压缩视觉信息,并将层次化信息从编码传递到解码。VisCo is a training-efficient self-compression framework that reuses the pretrained VLM itself as an intrinsic compressor that compresses visual information using a small set of memory tokens and transfers hierarchical information from encoding to decoding.

Spectral Prior for Reducing Exposure Bias in Diffusion Models
[标题中文] 降低扩散模型 Exposure Bias 的频谱先验
arXiv:2607.22091 多模态 方法 被引 0 · S2

提出 Spectral Alignment,一种基于 guidance 的轻量级方法,将中间预测的功率谱校准到预先计算的先验,并与 Classifier-Free Guidance (CFG) 互补。Spectral Alignment is proposed, a lightweight, guidance-based method that calibrates the power spectrum of intermediate predictions to a pre-computed prior and is complementary to Classifier-Free Guidance (CFG).

Multimodal Speaker Verification as a Threat to Speaker Anonymization
[标题中文] 多模态说话人验证对说话人匿名化的威胁
arXiv:2607.19636 多模态 方法 被引 0 · S2

该工作在多语句、多模态设定下研究 ASV,考察跨匿名语音聚合信息是否影响隐私,并发现帧级聚合得到的 EER 最低。This work investigates ASV in a multi-utterance, multimodal setting and examines whether aggregating information across anonymized speech impacts privacy, and finds that frame-level aggregation yields the lowest EERs.

Three-Body Scattering for Generative Modeling
[标题中文] 用于生成建模的三体散射
arXiv:2607.18198 多模态 方法 被引 1 · S2

这些结果将 tracked scattering 确立为通往高维 one-step generation 的路径,并给出一张设计图,将 diffusion 相关监督、Drift-like 动力学与 GAN-like 目标联系起来。These results establish tracked scattering as a route to high-dimensional one-step generation and provide a design map relating diffusion-related supervision, Drift-like dynamics, and GAN-like objectives.

O-VAD: Industrial Video Anomaly Detection through Object-Centric Tracking and Reasoning
O-VAD: 基于以对象为中心的跟踪与推理的工业视频异常检测
arXiv:2607.18142 多模态 方法 被引 0 · S2

该工作提出一个面向异常检测的免训练 agentic 框架,无需领域特定知识,旨在追踪被检测对象随时间变化的空间-时间动态与底层变换,然后基于逐对象的时间状态轨迹进行推理,在 grounding 帧中识别异常对象。This work introduces a training-free agentic framework for anomaly detection free of domain-specific knowledge, designed to track spatial-temporal dynamics and underlying transformations of detected objects over time, and then reason over the object-wise temporal state trajectories to identify abnormal objects in grounded frames.

Rethinking Classifier-Free Guidance in On-Policy Diffusion Distillation
重新思考 On-Policy 扩散蒸馏中的无分类器引导
arXiv:2607.24731 多模态 观点 OA · 绿色 被引 0 · S2 + OpenAlex

将 Positive--Direction Matching (PDM)——一种分支感知的 OPD 目标,分别约束正预测方向与 CFG 条件方向——引入 dense-to-sparse 视频控制;由于朴素的 guided matching 对推理 guidance 尺度极为敏感,分支感知监督可实现更鲁棒、更有效的知识迁移。Positive--Direction Matching (PDM), a branch-aware OPD objective that separately constrains the positive prediction and the CFG conditional direction, is introduced to dense-to-sparse video control, where naive guided matching is highly sensitive to inference guidance scales, while branch-aware supervision enables more robust and effective knowledge transfer.

Evidence Attribution in Visual Document Understanding without Coordinates or Region Labels
无坐标与区域标签的视觉文档理解中的证据归因
arXiv:2607.24651 多模态 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

本研究考察是否存在一条实用路径,在没有坐标界面、且无需高成本区域级监督的条件下提升归因效果,并指出了这样一条可行路径。A study investigates whether there is a practical path to improve attribution without a coordinate interface and without costly region-level supervision, and indicates a practical path to improve attribution without a coordinate interface and without costly region-level supervision.

Sol-Attn: Accelerating Video Generation Inference via On-the-Fly Attention Sparsification
Sol-Attn:通过即时注意力稀疏化加速视频生成推理
arXiv:2607.24027 多模态 方法 OA · 绿色 被引 2 · S2

本文提出无需训练的 Sol-Attn(Sparsifying online attention),在单次 online-softmax pass 中统一动态路由、稀疏计算与近似修正,在稀疏注意力中取得更好的精度–效率权衡。This paper introduces training-free Sol-Attn (Sparsifying online attention), which unifies dynamic routing, sparse computation, and approximation correction in a single online-softmax pass, achieving a better accuracy-efficiency trade-off in sparse attention.

Chamaileon: Cross-Context Binder Design with Contextualized Modeling and Mixed Sampling
Chamaileon:基于上下文建模与混合采样的跨上下文结合子设计
arXiv:2607.23518 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 Chamaileon,通过将问题建模为跨上下文结合景观(cross-context binding landscape modeling),统一多目标与多态 binder 设计,有效生成可适配多样构象景观与多目标需求的序列。Chamaileon is introduced, which unifies multi-target and multi-state binder design by formulating the problem as cross-context binding landscape modeling and effectively generates sequences adaptable to diverse conformational landscapes and multi-target requirements.

WorldDiT: A Unified Diffusion Architecture for World and Action Modeling
WorldDiT:面向世界建模与动作建模的统一扩散架构
arXiv:2607.23909 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

WorldDiT 是一种统一的 diffusion Transformer 架构,将动作生成与视觉世界建模耦合,无需大型预训练 VLM 动作主干即取得强性能,在报告全部四个 suite 的方法中,其总模型参数量与平均成功率处于已报告的 Pareto 前沿上。WorldDiT, a unified diffusion transformer architecture that couples action generation with visual world modeling and achieves strong performance without a large pretrained VLM action backbone, lies on the reported Pareto frontier for total model parameters and mean success among methods reporting all four suites.

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models
UltraViT:面向大视觉-语言模型的端侧延迟优化视觉编码器
arXiv:2607.23373 多模态 应用落地 OA · 绿色 被引 0 · S2 + OpenAlex

大量实验表明,结合端上延迟感知设计与定制化训练策略,建立了高效 LVLM 编码的新 SOTA,在端上以近 1.7 倍速度运行的同时显著优于现有以编码器为中心的基线。Extensive experiments demonstrate that the on-device latency-informed design combined with the tailored training strategy establishes a new state-of-the-art for efficient LVLM encoding, significantly outperforming existing encoder-centric baselines while operating on-device at nearly 1.7xthe speed.

Visual prompt engineering for video models
视频模型的视觉提示工程
arXiv:2607.25537 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文发现视觉提示工程(visual prompt engineering,简称 VIPE)能在多项任务上提升视频推理性能,甚至比经典的文本提示工程或 test-time scaling 更有效。It is found that visual prompt engineering, or VIPE for short, improves video reasoning performance across tasks and can be even more effective than classic text-based prompt engineering or test-time scaling.

Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model
Mage-VL:面向高效编解码器原生流式多模态的基础模型
arXiv:2607.24904 多模态 方法 OA · 绿色 被引 1 · S2

本文提出 Mage-VL,一种面向实时多模态理解与交互的高效 codec-native 流式基础模型,并构建了 AI4AI 数据流水线,涵盖面向多模态 captioning 的 prompt-code 联合优化与以 AI 驱动的性能诊断,以指导训练方案。Mage-VL is presented, an efficient codec-native streaming foundation model for real-time multimodal understanding and interaction and establishes AI4AI data pipelines encompassing prompt-code joint optimization for multimodal captioning and AI-driven performance diagnosis to guide training recipes.

Parallel Decoding Distillation for Fast Image and Video Generation
并行解码蒸馏:面向快速图像与视频生成
arXiv:2607.26004 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

介绍 Parallel Decoding Distillation,一种简化且可扩展的基于轨迹的蒸馏方法,用于 diffusion 和 flow matching 模型的快速推理,并显著提升生成视频的多样性。Parallel Decoding Distillation is introduced, a simplified and scalable trajectory-based distillation method for fast inference of diffusion and flow matching models and presents a significant improvement in generated video diversity.

Vision Mamba: Efficient Visual Representation Learning with Bidirectional State Space Model
Vision Mamba:基于双向状态空间模型的高效视觉表示学习
arXiv:2401.09417 多模态 方法 OA · 绿色 被引 2083 · S2

本文提出基于双向 Mamba 块(Vim)的通用视觉 backbone,通过位置嵌入标记图像序列,并利用双向 state space model 压缩视觉表征,具有成为下一代视觉基础模型 backbone 的巨大潜力。This paper proposes a new generic vision backbone with bidirectional Mamba blocks (Vim), which marks the image sequences with position embeddings and compresses the visual representation with bidirectional state space models and has great potential to be the next-generation backbone for vision foundation models.

PaLM-E: An Embodied Multimodal Language Model
PaLM-E:一种具身多模态语言模型
arXiv:2303.03378 多模态 方法 OA · 绿色 被引 3032 · S2

本文提出 embodied language model,将真实世界连续传感器模态直接融入语言模型,从而建立词语与感知之间的联系,实现真实世界中的通用推理。This work proposes embodied language models to directly incorporate real-world continuous sensor modalities into language models and thereby establish the link between words and percepts to enable general inference in the real world.

SimVLM: Simple Visual Language Model Pretraining with Weak Supervision
SimVLM:基于弱监督的简单视觉语言模型预训练
arXiv:2108.10904 多模态 方法 OA · 绿色 被引 970 · S2

本文提出极简的预训练框架 SimVLM,在广泛的判别式与生成式视觉-语言基准上显著超越既往预训练方法并取得新 SOTA,包括 VQA、NLVR2 以及图像描述任务。This work presents a minimalist pretraining framework, named SimVLM, which significantly outperforms previous pretraining methods and achieves new state-of-the-art results on a wide range of discriminative and generative vision-language benchmarks, including VQA, NLVR2, and image captioning tasks.

Human-in-the-Loop Signature Bootstrapping for UAV Hyperspectral PFM-1 Mine Detection
UAV 高光谱 PFM-1 地雷检测中的人在回路签名引导
arXiv:2607.25310 多模态 应用落地 OA · 绿色 被引 1 · S2

本文研究无人机(UAV)可见光-近红外(VNIR)高光谱图像中 PFM-1 地雷的检测,使用光谱角制图(SAM)、匹配滤波器(MF)、自适应相干估计器(ACE)和约束能量最小化(CEM)。This paper studies PFM-1 landmine detection in unmanned aerial vehicle (UAV) visible and near-infrared (VNIR) HSI using spectral angle mapper (SAM), matched filter (MF), adaptive coherence estimator (ACE), and constrained energy minimization (CEM).

CLBench-V: Evaluating Multimodal Context Learning from Grounding to Knowledge Acquisition
CLBench-V: 评估多模态上下文学习——从 grounding 到知识获取
arXiv:2607.25294 多模态 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

本文介绍 CLBench-V,一个多模态上下文学习 benchmark,围绕三个维度组织任务——上下文 grounding、新信息应用与新知识学习——以解决定位上下文使用失效位置的难题。This work introduces CLBench-V, a benchmark for multimodal context learning that addresses the difficulty of localizing where context use breaks down by organizing tasks around three dimensions: context grounding, new information application, and new knowledge learning.

Understanding Is Done Early: A Depth Division of Labor in Large Language Models and Its Use for Unbounded-Context Memory
理解在前部完成:大语言模型中的深度分工及其在无界上下文记忆中的应用
arXiv:2607.28263 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

结果表明,长上下文 memory 可沿 layer 轴(而非仅沿 token 轴)进行组织,并揭示了有界检索的优势及其在窗口内的压缩代价。These results show that long-context memory can be organized along the layer axis, not only the token axis, and expose both the benefits of bounded retrieval and its in-window compression tax.

ShadowDancer: Teaching Video World Models Any Action by Learning Unified Dynamics Representations from a Video and Its Shadow
ShadowDancer: 通过从视频及其阴影中学习统一动力学表示来教视频世界模型执行任意动作
arXiv:2607.28362 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

ShadowDancer 引入 shadow pair,即在同一动力学下对外观做独立重采样的成对视频,并由 Shadow Library 大规模构建;一个 dynamics family 可控,当且仅当能为其构造出这样的 pair。ShadowDancer introduces shadow pairs, video pairs that replay the same dynamics under independently resampled appearance, constructed at scale by the Shadow Library, so that a dynamics family becomes controllable exactly when such pairs can be constructed for it.

See2Think: Do Multimodal Models Really Use Intermediate Visual States?
See2Think:多模态模型真的使用了中间视觉状态吗?
arXiv:2607.26769 多模态 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

对代表性闭源与开源多模态模型的评测表明,视觉推理强依赖于模型与环境,没有任何单一设置能在所有任务上持续占优。Evaluating representative proprietary and open-source multimodal models, it is found that visual reasoning is strongly model- and environment-dependent, with no single setting consistently dominating across tasks.