Papers · organized/paper_cards

论文

156 张论文卡片

开放获取 全部 绿色 · 724
Vision Mamba: Efficient Visual Representation Learning with Bidirectional State Space Model
Vision Mamba:基于双向状态空间模型的高效视觉表示学习
arXiv:2401.09417 多模态 方法 OA · 绿色 被引 2083 · S2

本文提出基于双向 Mamba 块(Vim)的通用视觉 backbone,通过位置嵌入标记图像序列,并利用双向 state space model 压缩视觉表征,具有成为下一代视觉基础模型 backbone 的巨大潜力。This paper proposes a new generic vision backbone with bidirectional Mamba blocks (Vim), which marks the image sequences with position embeddings and compresses the visual representation with bidirectional state space models and has great potential to be the next-generation backbone for vision foundation models.

PaLM-E: An Embodied Multimodal Language Model
PaLM-E:一种具身多模态语言模型
arXiv:2303.03378 多模态 方法 OA · 绿色 被引 3032 · S2

本文提出 embodied language model,将真实世界连续传感器模态直接融入语言模型,从而建立词语与感知之间的联系,实现真实世界中的通用推理。This work proposes embodied language models to directly incorporate real-world continuous sensor modalities into language models and thereby establish the link between words and percepts to enable general inference in the real world.

SimVLM: Simple Visual Language Model Pretraining with Weak Supervision
SimVLM:基于弱监督的简单视觉语言模型预训练
arXiv:2108.10904 多模态 方法 OA · 绿色 被引 970 · S2

本文提出极简的预训练框架 SimVLM,在广泛的判别式与生成式视觉-语言基准上显著超越既往预训练方法并取得新 SOTA,包括 VQA、NLVR2 以及图像描述任务。This work presents a minimalist pretraining framework, named SimVLM, which significantly outperforms previous pretraining methods and achieves new state-of-the-art results on a wide range of discriminative and generative vision-language benchmarks, including VQA, NLVR2, and image captioning tasks.

VisualBERT: A Simple and Performant Baseline for Vision and Language
VisualBERT: 一个简单且性能优异的视觉与语言基线模型
arXiv:1908.03557 多模态 方法 OA · 绿色 被引 2375 · S2

分析表明,VisualBERT 能够在无任何显式监督的情况下将语言元素对应到图像区域,甚至对句法关系敏感,例如能够跟踪动词与其论元对应图像区域之间的关联。Analysis demonstrates that VisualBERT can ground elements of language to image regions without any explicit supervision and is even sensitive to syntactic relationships, tracking, for example, associations between verbs and image regions corresponding to their arguments.

BEiT: BERT Pre-Training of Image Transformers
BEiT: 图像 Transformers 的 BERT 预训练
arXiv:2106.08254 工程化 方法 OA · 绿色 被引 3854 · S2

文章介绍了一种自监督视觉表征模型 BEiT(Bidirectional Encoder representation from Image Transformers),在图像分类和语义分割上的结果表明,该模型取得了与先前预训练方法相当的竞争性结果。A self-supervised vision representation model BEiT, which stands for Bidirectional Encoder representation from Image Transformers, is introduced, and results on image classification and semantic segmentation show that the model achieves competitive results with previous pre-training methods.

Swin-Unet: Unet-like Pure Transformer for Medical Image Segmentation
Swin-Unet: 类 U-Net 的纯 Transformer 医学图像分割网络
arXiv:2105.05537 多模态 方法 OA · 绿色 被引 5572 · S2

在输入和输出直接进行 4 倍下采样与上采样的设定下,实验表明,基于纯 Transformer 的 U 形编码器-解码器网络优于完全卷积或 Transformer 与卷积相结合的方法。Under the direct down-sampling and up-sampled of the inputs and outputs by 4x, experiments demonstrate that the pure Transformer-based U-shaped Encoder-Decoder network outperforms those methods with full Convolution or the combination of transformer and convolution.

BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation
BLIP: 面向统一视觉-语言理解与生成的引导式语言-图像预训练
arXiv:2201.12086 多模态 方法 OA · 绿色 被引 7222 · S2

BLIP 通过引导式 caption 方式有效利用含噪网络数据,由 captioner 生成合成 caption,并由 filter 去除噪声样本;在以零样本方式直接迁移到视频-语言任务时,展现出强大的泛化能力。BLIP effectively utilizes the noisy web data by bootstrapping the captions, where a captioner generates synthetic captions and a filter removes the noisy ones, and demonstrates strong generalization ability when directly transferred to video-language tasks in a zero-shot manner.

A deep learning architecture for temporal sleep stage classification\n using multivariate and multimodal time series
基于多变量多模态时间序列的深度学习睡眠阶段时序分类架构
arXiv:1707.03321 多模态 方法 OA · 绿色 被引 615 · OpenAlex

本文提出了首个用于睡眠阶段分类的深度学习方法,无需计算频谱图或提取手工特征即可端到端学习,利用了全部多变量多模态 PSG 信号(EEG、EMG、EOG),并能利用每个 30 秒窗口数据的时序上下文。This work introduces here the first deep learning approach for sleep stage classification that learns end-to-end without computing spectrograms or extracting handcrafted features, that exploits all multivariate and multimodal polysomnography (PSG) signals (EEG, EMG, and EOG), and that can exploit the temporal context of each 30-s window of data.

Multi-Task Learning Using Uncertainty to Weigh Losses for Scene Geometry and Semantics
利用不确定性为损失加权的多任务学习,用于场景几何与语义
arXiv:1705.07115 工程化 方法 OA · 绿色 被引 4348 · S2

本文提出一种多任务深度学习的原则性方法,通过考虑各任务的同方差不确定性来加权多个损失函数,从而在分类与回归场景下同时学习具有不同单位或尺度的多种量。A principled approach to multi-task deep learning is proposed which weighs multiple loss functions by considering the homoscedastic uncertainty of each task, allowing us to simultaneously learn various quantities with different units or scales in both classification and regression settings.

Deep Fragment Embeddings for Bidirectional Image Sentence Mapping
用于双向图文映射的深度片段嵌入
arXiv:1406.5679 RAG 检索增强 方法 OA · 绿色 被引 982 · S2

本文通过深度多模态嵌入视觉与自然语言数据,提出了一种用于图文双向检索的模型,并引入结构化的最大间隔目标,使该模型能够显式地跨模态关联片段。This work introduces a model for bidirectional retrieval of images and sentences through a deep, multi-modal embedding of visual and natural language data and introduces a structured max-margin objective that allows this model to explicitly associate fragments across modalities.

A Survey on Multi-Task Learning
多任务学习综述
arXiv:1707.08114 工程化 综述 OA · 绿色 被引 3073 · S2

从算法建模、应用和理论分析角度对 MTL 的综述,给出了 MTL 的定义,并将不同 MTL 算法分为五类:特征学习方法、低秩方法、任务聚类方法、任务关系学习方法和分解方法A survey for MTL from the perspective of algorithmic modeling, applications and theoretical analyses, which gives a definition of MTL and classify different MTL algorithms into five categories, including feature learning approach, low-rank approach, task clustering approach,task relation learning approach and decomposition approach.

Skip-Thought Vectors
Skip-Thought Vectors
arXiv:1506.06726 LLM 基础设施 方法 OA · 绿色 被引 2488 · S2

描述了一种无监督学习通用分布式句子编码器的方法,利用书籍文本的连续性,训练编码器-解码器模型以重建编码段落的周围句子。The approach for unsupervised learning of a generic, distributed sentence encoder is described, using the continuity of text from books to train an encoder-decoder model that tries to reconstruct the surrounding sentences of an encoded passage.