Papers · organized/paper_cards

论文

258 张论文卡片 · 多模态

开放获取 全部 绿色 · 724
DINOv2: Learning Robust Visual Features without Supervision
DINOv2:无监督学习鲁棒的视觉特征
arXiv:2304.07193 多模态 方法 OA · 绿色 被引 9976 · S2

本文回顾现有方法,并融合多种技术从数据与模型规模两方面扩展预训练,提出一条自动化流水线以构建专用、多样且经过筛选的图像数据集,替代自监督文献中常用的未筛选数据。This work revisits existing approaches and combines different techniques to scale the pretraining in terms of data and model size, and proposes an automatic pipeline to build a dedicated, diverse, and curated image dataset instead of uncurated data, as typically done in the self-supervised literature.

ScanNet: Richly-annotated 3D Reconstructions of Indoor Scenes
ScanNet:富含标注的室内场景三维重建
arXiv:1702.04405 多模态 评测集 OA · 绿色 被引 5793 · S2

本文推出 ScanNet,一个 RGB-D 视频数据集,包含 1513 个场景中的 250 万视角,标注有三维相机位姿、表面重建与语义分割,并表明使用该数据可在多项三维场景理解任务上取得 SOTA 性能。This work introduces ScanNet, an RGB-D video dataset containing 2.5M views in 1513 scenes annotated with 3D camera poses, surface reconstructions, and semantic segmentations, and shows that using this data helps achieve state-of-the-art performance on several 3D scene understanding tasks.

A Multitask, Multilingual, Multimodal Evaluation of ChatGPT on Reasoning, Hallucination, and Interactivity
ChatGPT 在推理、幻觉与交互性方面的多任务、多语言、多模态评估
arXiv:2302.04023 多模态 评测集 OA · 绿色 被引 1809 · S2

研究发现 ChatGPT 在大多数任务上以零样本学习优于其他 LLM,在部分任务上甚至超过微调模型,并且对非拉丁文字语言的理解能力优于生成能力。It is found that ChatGPT outperforms LLMs with zero-shot learning on most tasks and even outperforms fine-tuned models on some tasks and is better at understanding non-Latin script languages than generating them.

Matterport3D: Learning from RGB-D Data in Indoor Environments
Matterport3D:基于室内 RGB-D 数据的学习
arXiv:1709.06158 多模态 评测集 OA · 绿色 被引 2631 · S2

本文介绍 Matterport3D,一个大规模 RGB-D 数据集,包含来自 90 个建筑物级场景共 194,400 张 RGB-D 图像的 10,800 个全景视图,可支持多种监督与自监督计算机视觉任务,包括关键点匹配、视角重叠预测、由彩色图像预测法线、语义分割和区域分类。Matterport3D is introduced, a large-scale RGB-D dataset containing 10,800 panoramic views from 194,400RGB-D images of 90 building-scale scenes that enable a variety of supervised and self-supervised computer vision tasks, including keypoint matching, view overlap prediction, normal prediction from color, semantic segmentation, and region classification.

Embracing Imperfect Datasets: A Review of Deep Learning Solutions for Medical Image Segmentation
拥抱不完美数据集:医学图像分割中深度学习解决方案综述
arXiv:1908.10454 多模态 综述 OA · 绿色 被引 1035 · S2

本文对上述解决方案进行了详细综述,总结了其技术创新与实验结果,比较了各方法的优势与适用条件,并给出推荐方案。This article provides a detailed review of the solutions above, summarizing both the technical novelties and empirical results, and compares the benefits and requirements of the surveyed methodologies and provides recommended solutions.

Recent Advances in Convolutional Neural Networks
卷积神经网络近期进展
arXiv:1512.07108 多模态 综述 OA · 绿色 被引 6068 · S2

本文详细介绍了 CNN 在多个方面的改进,包括层设计、激活函数、损失函数、正则化、优化与快速计算,并阐述了卷积神经网络在计算机视觉、语音与自然语言处理中的多种应用。This paper details the improvements of CNN on different aspects, including layer design, activation function, loss function, regularization, optimization and fast computation, and introduces various applications of convolutional neural networks in computer vision, speech and natural language processing.

See like a Robot: Robot-Centric Pointmaps for Vision-Language-Action Models
像机器人一样看:面向视觉-语言-动作模型的机器人中心点图
arXiv:2607.11498 多模态 方法 被引 0 · S2

Pointmaps 在保留预训练 2D VLA 所需 H × W 稠密网格的同时,提供机器人坐标系下的 3D 几何信息,能以极小的架构改动集成到现有 VLA 中,并提升 pi0.5 与 SmolVLA 的性能,优于代表性的相机视点和 3D 感知基线。Pointmaps provide robot-frame 3D geometry while preserving the dense H x W grid expected by pretrained 2D VLAs, so they integrate into existing VLAs with minimal architectural change and improve both pi0.5 and SmolVLA and outperform representative camera-viewpoint and 3D-aware baselines.

SVR-R1: Bootstrapping Multi-modal Reasoning with Self-verification in Reinforcement Learning
SVR-R1:在强化学习中通过自验证自举多模态推理
arXiv:2607.10966 多模态 评测集 被引 0 · S2

本文提出自验证推理器 SVR-R1,一种多轮强化学习框架,将模型自身的验证转化为多模态推理的学习信号,提供了一种简洁而有效的多模态推理自举方案。Self-Verified Reasoner (SVR-R1), a multi-turn RL framework that turns a model's own verification into a learning signal for multimodal reasoning, is introduced, offering a simple yet effective recipe for bootstrapping multimodal reasoning.

HOMIE: Human-object Centric Video Personalization via Multimodal Intelligent Enchancement
HOMIE:通过多模态智能增强实现以人-物为中心的视频个性化
arXiv:2607.18217 多模态 方法 被引 0 · S2

HOMIE 提出了一种更优的 MLLM 集成策略,可在不损害文本编码器可控性或引入昂贵重新对齐的前提下,提取参考级关系知识,并在 self-attention 中引入全局多模态引导,使 MLLM 派生的语义特征与 VAE token 更好对齐。HOMIE proposes a better MLLM integration strategy to extract knowledge of reference-level relationships without compromising the controllability of text encoders or incurring costly re-alignment, and introduces global multimodal guidance within self-attention to better align MLLM-derived semantic features with VAE tokens.

TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs
TimeLens2:基于多模态 LLM 的通用视频时序定位
arXiv:2607.17423 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

在七个基准上,TimeLens2-2B 在所有基准上均优于规模相当的所有基线,4B 和 8B 变体则取得了 SOTA 性能,超越了参数量高达 397B 的开源模型。Across seven benchmarks, TimeLens2-2B outperforms all size-matched baselines on every benchmark, while the 4B and 8B variants achieve state-of-the-art performance, surpassing open-source models with up to 397B parameters.

ReflectWorld-MM: An Entity-Oriented Multimodal Memory System for Open-Ended Video Streams
ReflectWorld-MM:面向开放视频流的实体导向多模态记忆系统
arXiv:2607.09759 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出了ReflectWorld-MM,一个面向开放视频流的以实体为中心的多模态记忆系统,在六个长视频和终身记忆基准测试上均达到最优准确率,超越了强记忆Agent和前沿模型。ReflectWorld-MM is proposed, an entity-oriented multimodal memory system for open-ended video streams that achieves the best accuracy on all six long-video and lifelong-memory benchmarks, outperforming strong memory agents and a frontier model.

ShotPlan: Cinematic Video Generation with Learnable Planning Token
ShotPlan:基于可学习规划token的电影级视频生成
arXiv:2607.17675 多模态 观点 OA · 绿色 被引 0 · S2 + OpenAlex

提出了ShotPlan,一个基于视频扩散基础模型构建的、用于显式多镜头电影级视频生成的框架,显著优于现有的电影级视频生成方法,提供更灵活的镜头管理和更强的跨镜头一致性。ShotPlan is proposed, a framework for explicit multi-shot cinematic video generation built upon a video diffusion foundation model that significantly outperforms existing cinematic video generation methods, offering more flexible shot management and stronger inter-shot consistency.

Can Multimodal Large Language Models Understand OCT?
多模态大语言模型能理解OCT吗?
arXiv:2607.16609 多模态 评测集 OA · 绿色 被引 2 · S2

OCT-Bench能够对MLLM进行全面且细粒度的评估,为识别能力瓶颈和推进临床可信的OCT理解奠定基础。OCT-Bench enables comprehensive and fine-grained evaluation of MLLMs, providing a foundation for identifying capability bottlenecks and advancing clinically grounded OCT understanding.

GigaChat Audio: Time-aware Large Audio Language Model
GigaChat Audio:时间感知的大音频语言模型
arXiv:2607.10387 多模态 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出一个时间感知的音频 LLM,能够基于大规模合成监督(来自级联 pipeline)在长达 120 分钟的输入上回答带有显式时间戳的问题,并在短时长和长时长 benchmark 上取得强劲的时间定位准确率。This work presents a time-aware audio LLM that answers questions with explicit timestamps over up to 120 minutes of input using large-scale synthetic supervision from a cascaded pipeline and achieves strong temporal-grounding accuracy on short and long benchmarks.

Very Deep Convolutional Networks for Large-Scale Image Recognition
Very Deep Convolutional Networks for Large-Scale Image Recognition
arXiv:1409.1556 多模态 方法 OA · 绿色 被引 113645 · S2

本文研究了在采用极小卷积滤波器的架构下,卷积网络深度对大规模图像识别精度的影响,并表明将深度推进至 16-19 个权重层,可在先前 SOTA 配置基础上取得显著提升。This work investigates the effect of the convolutional network depth on its accuracy in the large-scale image recognition setting using an architecture with very small convolution filters, which shows that a significant improvement on the prior-art configurations can be achieved by pushing the depth to 16-19 weight layers.

VQA: Visual Question Answering
VQA: Visual Question Answering
arXiv:1505.00468 多模态 评测集 OA · 绿色 被引 6654 · S2
HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million\n Narrated Video Clips
HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video Clips
arXiv:1906.03327 多模态 方法 OA · 绿色 被引 1510 · S2

在 YouCook2、CrossTask 等教学视频数据集上,基于该数据训练的文本-视频 embedding 在文本到视频检索与动作定位任务上达到了 SOTA 结果。It is demonstrated that a text-video embedding trained on this data leads to state-of-the-art results for text-to-video retrieval and action localization on instructional video datasets such as YouCook2 or CrossTask.

Align before Fuse: Vision and Language Representation Learning with Momentum Distillation
Align before Fuse: Vision and Language Representation Learning with Momentum Distillation
arXiv:2107.07651 多模态 方法 OA · 绿色 被引 2913 · S2

提出在通过跨模态注意力融合之前对齐图像与文本表示的对比损失(ALBEF),可实现更扎实的视觉-语言表征学习;并提出动量蒸馏,一种利用动量模型生成伪目标进行自训练的方法。A contrastive loss to ALign the image and text representations BEfore Fusing (ALBEF) them through cross-modal attention, which enables more grounded vision and language representation learning and proposes momentum distillation, a self-training method which learns from pseudo-targets produced by a momentum model.

Visual Instruction Tuning
Visual Instruction Tuning
arXiv:2304.08485 多模态 方法 OA · 绿色 被引 10838 · S2

本文提出 LLaVA:Large Language and Vision Assistant,一个端到端训练的大型多模态模型,将视觉编码器与 LLM 相结合用于通用视觉和语言理解;并引入 GPT-4 生成的视觉指令微调数据,模型与代码库已开源。This paper presents LLaVA: Large Language and Vision Assistant, an end-to-end trained large multimodal model that connects a vision encoder and LLM for general-purpose visual and language understanding and introduces GPT-4 generated visual instruction tuning data, the model and code base publicly available.

Deep Captioning with Multimodal Recurrent Neural Networks (m-RNN)
Deep Captioning with Multimodal Recurrent Neural Networks (m-RNN)
arXiv:1412.6632 多模态 方法 OA · 绿色 被引 1283 · S2

m-RNN 模型直接对给定先前词语和图像条件下生成下一个词的概率分布建模,相较于直接优化排序目标函数进行检索的 SOTA 方法,取得了显著的性能提升。The m-RNN model directly models the probability distribution of generating a word given previous words and an image, and achieves significant performance improvement over the state-of-the-art methods which directly optimize the ranking objective function for retrieval.

ViLT: Vision-and-Language Transformer Without Convolution or Region Supervision
ViLT: Vision-and-Language Transformer Without Convolution or Region Supervision
arXiv:2102.03334 多模态 方法 OA · 绿色 被引 2389 · S2

提出极简的 VLP 模型 Vision-and-Language Transformer (ViLT),其一体化设计将视觉输入处理大幅简化为与文本输入相同的无卷积方式;ViLT 比此前的 VLP 模型快达数十倍,同时下游任务性能具有竞争力甚至更优。A minimal VLP model, Vision-and-Language Transformer (ViLT), monolithic in the sense that the processing of visual inputs is drastically simplified to just the same convolution-free manner that the authors process textual inputs, showing that ViLT is up to tens of times faster than previous VLP models, yet with competitive or better downstream task performance.

CoCa: Contrastive Captioners are Image-Text Foundation Models
CoCa: Contrastive Captioners are Image-Text Foundation Models
arXiv:2205.01917 多模态 方法 OA · 绿色 被引 1813 · S2

Contrastive Captioner (CoCa) 采用极简设计,对图文编码器-解码器基础模型联合使用对比损失与字幕损失进行预训练,从而兼具 CLIP 等对比方法与 SimVLM 等生成方法的能力。Contrastive Captioner (CoCa), a minimalist design to pretrain an image-text encoder-decoder foundation model jointly with contrastive loss and captioning loss, thereby subsuming model capabilities from contrastive approaches like CLIP and generative methods like SimVLM.

MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models
MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models
arXiv:2304.10592 多模态 方法 OA · 绿色 被引 3316 · S2

本文提出 MiniGPT-4,通过一个投影层将冻结的视觉编码器与冻结的先进 LLM Vicuna 对齐,发现将视觉特征与先进大语言模型恰当对齐可获得类似 GPT-4 所展现的多种先进多模态能力。MiniGPT-4 is presented, which aligns a frozen visual encoder with a frozen advanced LLM, Vicuna, using one projection layer to uncovers that properly aligning the visual features with an advanced large language model can possess numerous advanced multi-modal abilities demonstrated by G PT-4.

VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training
VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training
arXiv:2203.12602 多模态 方法 OA · 绿色 被引 2186 · S2

本文表明视频掩码自编码器(VideoMAE)是自监督视频预训练(SSVP)的数据高效学习器,并受近期 ImageMAE 启发,提出采用极高掩码比例的定制化视频管状掩码策略。This paper shows that video masked autoencoders (VideoMAE) are data-efficient learners for self-supervised video pre-training (SSVP), and proposes customized video tube masking with an extremely high ratio, inspired by the recent ImageMAE.

InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning
InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning
arXiv:2305.06500 多模态 方法 OA · 绿色 被引 3878 · S2

本文基于预训练 BLIP-2 模型,对视觉-语言指令微调展开系统全面研究,并提出指令感知的 Query Transformer,用于提取针对给定指令的信息丰富特征。This paper conducts a systematic and comprehensive study on vision-language instruction tuning based on the pretrained BLIP-2 models, and introduces an instruction-aware Query Transformer, which extracts informative features tailored to the given instruction.

ReViV: Reconstructing the Viewer and the View in 4D from Monocular Egocentric Video
ReViV:从单目自我中心视频 4D 重建观察者与视角
arXiv:2607.17790 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

ReViV 是首个用于整体第一人称 4D 重建的统一框架,能够从单个单目 RGB 视频中同时提取观察者与视角动态,在整体自身体、手部与注视重建以及相机跟踪方面达到 SOTA 精度与效率,同时保持极具竞争力的第一人称深度估计能力。ReViV is the first unified framework for holistic egocentric 4D reconstruction that extracts both viewer and view dynamics from a single monocular RGB video and achieves state-of-the-art accuracy and efficiency across holistic ego-body, hand, and gaze reconstruction, camera tracking, while maintaining highly competitive egocentric depth estimation.

ABot-World-0: Infinite Interactive World Rollout on a Single Desktop GPU
ABot-World-0:在单台桌面级 GPU 上实现无限交互式世界推演
arXiv:2607.19191 多模态 方法 OA · 绿色 被引 1 · S2

本工作通过 teacher forcing 与 ODE 蒸馏,将一个双向动作条件教师模型逐步蒸馏为因果学生模型,并提出 LongForcing,将学生模型的长时间自展开与扩展时域教师模型对齐,从而缓解累积的分布漂移与自回归漂移。This work progressively distill a bidirectional action-conditioned teacher into a causal student through teacher forcing and ODE distillation, and introduces LongForcing to align long student self-rollouts with an extended-horizon teacher, mitigating accumulated distribution shift and autoregressive drift.

Generative Adversarial Networks
生成对抗网络
arXiv:1406.2661 多模态 方法 OA · 绿色 被引 6813 · S2
Momentum Contrast for Unsupervised Visual Representation Learning
无监督视觉表征学习的动量对比
arXiv:1911.05722 多模态 方法 OA · 绿色 被引 15586 · S2
Object Detection in 20 Years: A Survey
目标检测二十年:综述
arXiv:1905.05055 多模态 综述 OA · 绿色 被引 3564 · S2

本文从技术演进的角度,对这一快速发展的研究领域进行了广泛综述,跨越超过四分之一世纪的时间跨度(从 1990 年代到 2022 年)。This article extensively reviews this fast-moving research field in the light of technical evolution, spanning over a quarter-century’s time (from the 1990s to 2022).

Computational Humor with Multimodal LLMs: Methods, Datasets, Evaluation, and Challenges
多模态 LLM 的计算幽默:方法、数据集、评估与挑战
arXiv:2607.19011 多模态 综述 OA · 绿色 被引 0 · S2 + OpenAlex

本综述聚焦于单图与多格视觉作品中的幽默理解,同时将幽默生成视为新兴的下游前沿方向,并围绕多模态对齐、证据 grounded 推理与可控生成,对基准设计、评估协议与建模范式进行系统综述。This survey focuses on visual humor understanding in single-image and multi-panel artifacts, while treating humor generation as an emerging downstream frontier, and synthesizes benchmark design, evaluation protocols, and modeling paradigms based on multimodal alignment, evidence-grounded reasoning, and controlled generation.

Self Gradient Forcing: Native Long Video Extrapolation
Self Gradient Forcing:原生长视频外推
arXiv:2607.20368 多模态 方法 OA · 绿色 被引 1 · S2

Self Gradient Forcing(SGF)是一种两阶段训练策略,在原生自回归训练目标内恢复缺失的"记忆写入"监督信号,通过对未来视频 latent 的损失来训练模型将上下文编码为更有效的因果记忆。Self Gradient Forcing (SGF), a two-pass training strategy that restores this missing memory-writing supervision within the native autoregressive training objective, using losses on future video latents to train the model to encode context into more effective causal memory.

Trace: A Taxonomy-Guided Environment for Multidomain Visual Reasoning
Trace:面向多领域视觉推理的 Taxonomy 引导环境
arXiv:2607.19790 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 Trace,一个面向多领域视觉推理的、由分类体系引导的环境,其将任务构建分解为场景语法与可执行任务程序,将视觉呈现与答案计算解耦,并提供了广泛的程序化训练可迁移到生成任务分布之外的证据。Trace is introduced, a taxonomy-guided environment for multidomain visual reasoning that factorizes task construction into a scene grammar and an executable task program, separating visual realization from answer computation, providing evidence that broad procedural training can transfer beyond the generated task distributions.

An Exam for Active Observers
面向主动观察者的评测
arXiv:2607.16165 多模态 评测集 OA · 绿色 被引 0 · S2 + OpenAlex

人类视觉是一个闭环:注视点不断被中间假设而非单一快照持续重定向。数十年的心理物理学与认知科学研究表明,主动观察对多种任务至关重要。当代多模态大语言模型 (MLLM) 是否进行主动观察,是一个现有视觉语言基准无法回答的经验问题。我们提出 ActiveVision,一个使 MLLM 主动观察可度量的基准,包含 3 个类别共 17 个任务,任务设计强制进行重复视觉感知……Human vision is a closed loop: gaze is continuously redirected by intermediate hypotheses rather than a single snapshot. Decades of psychophysics and cognitive science have argued that this active observation is essential for a wide range of tasks. Whether today's multimodal large language models (MLLMs) exercise active observation is an empirical question that current vision-language benchmarks do not answer. We introduce ActiveVision, a benchmark that makes active observation measurable for MLLMs, comprising 17 tasks across 3 categories. Tasks are designed to force repeated visual perception

FVAttn: Adaptive Sparse Attention with Runtime Load Balancing for Video Generation
FVAttn:面向视频生成的自适应稀疏注意力与运行时负载均衡
arXiv:2607.16190 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

本文提出 \method,一种无需训练的稀疏注意力系统,可在多 GPU 序列并行下提升自适应稀疏注意力的分布式执行效率。This work presents \method, a training-free sparse-attention system that improves the distributed execution efficiency of adaptive sparse attention under multi-GPU sequence parallelism.

Train the Model, Not the Reader: Decodability Supervision for Verifiable Activation Explanations
训练模型而非读者:用于可验证激活解释的可解码性监督
arXiv:2607.20379 多模态 方法 OA · 绿色 被引 0 · S2 + OpenAlex

提出两项审计协议——grounding and truth 对比以及 swap to an independent evaluator,以及 RECAP(Readable Encodings via Co-trained Auxiliary Predictors),即与目标模型联合训练的线性头,用于保持指定内容的可解码性。Two audit protocols, the comparison of grounding and truth and the swap to an independent evaluator, and RECAP (Readable Encodings via Co-trained Auxiliary Predictors), linear heads trained alongside the target model to keep designated content decodable are contributed.