Papers · organized/paper_cards

论文

10 张论文卡片 · 多模态 · 综述

开放获取 全部 绿色 · 677
RedVox: Safety and Fairness Gaps in Speech Models Across Languages
RedVox:语音模型跨语言的安全性与公平性差距
arXiv:2606.26968 多模态 综述 OA · 绿色 被引 0 · S2 + OpenAlex

提出 RedVox,一个基于真实人声构建的音频与语音多语言安全性与公平性基准,涵盖五种语言中的不安全与不公平的刻板请求;研究发现漏洞即使在非对抗条件下仍然存在,在非英语语言中更为严重,且在请求来自语音输入时会被进一步放大。RedVox is introduced, a multilingual safety and fairness benchmark for audio and speech built on real voices, covering unsafe and unfair stereotypical requests across five languages, finding that vulnerabilities persist even under non-adversarial conditions, worsen in non-English languages, and are amplified when the request comes from a spoken input.

A Comprehensive Survey of AI-Generated Content (AIGC): A History of Generative AI from GAN to ChatGPT
A Comprehensive Survey of AI-Generated Content (AIGC): A History of Generative AI from GAN to ChatGPT
arXiv:2303.04226 多模态 综述 OA · 绿色 被引 838 · S2

该综述全面回顾了生成模型的历史与基本组件,以及 AIGC 在单模态交互与多模态交互方向的最新进展,并介绍了文本与图像生成任务及相关模型。This survey provides a comprehensive review on the history of generative models, and basic components, recent advances in AIGC from unimmodal interaction and multimodal interaction, and introduces the generation tasks and relative models of text and image.

A Review on Generative Adversarial Networks: Algorithms, Theory, and Applications
A Review on Generative Adversarial Networks: Algorithms, Theory, and Applications
arXiv:2001.06937 多模态 综述 OA · 绿色 被引 1154 · S2

详细介绍大多数 GAN 算法的动机、数学表示与结构,并对它们的共性与差异进行比较。The motivations, mathematical representations, and structures of most GAN algorithms are introduced in detail, and they are compared to compare their commonalities and differences.

Embracing Imperfect Datasets: A Review of Deep Learning Solutions for Medical Image Segmentation
拥抱不完美数据集:医学图像分割中深度学习解决方案综述
arXiv:1908.10454 多模态 综述 OA · 绿色 被引 1035 · S2

本文对上述解决方案进行了详细综述,总结了其技术创新与实验结果,比较了各方法的优势与适用条件,并给出推荐方案。This article provides a detailed review of the solutions above, summarizing both the technical novelties and empirical results, and compares the benefits and requirements of the surveyed methodologies and provides recommended solutions.

Recent Advances in Convolutional Neural Networks
卷积神经网络近期进展
arXiv:1512.07108 多模态 综述 OA · 绿色 被引 6068 · S2

本文详细介绍了 CNN 在多个方面的改进,包括层设计、激活函数、损失函数、正则化、优化与快速计算,并阐述了卷积神经网络在计算机视觉、语音与自然语言处理中的多种应用。This paper details the improvements of CNN on different aspects, including layer design, activation function, loss function, regularization, optimization and fast computation, and introduces various applications of convolutional neural networks in computer vision, speech and natural language processing.

Object Detection in 20 Years: A Survey
目标检测二十年:综述
arXiv:1905.05055 多模态 综述 OA · 绿色 被引 3564 · S2

本文从技术演进的角度,对这一快速发展的研究领域进行了广泛综述,跨越超过四分之一世纪的时间跨度(从 1990 年代到 2022 年)。This article extensively reviews this fast-moving research field in the light of technical evolution, spanning over a quarter-century’s time (from the 1990s to 2022).

Computational Humor with Multimodal LLMs: Methods, Datasets, Evaluation, and Challenges
多模态 LLM 的计算幽默:方法、数据集、评估与挑战
arXiv:2607.19011 多模态 综述 OA · 绿色 被引 0 · S2 + OpenAlex

本综述聚焦于单图与多格视觉作品中的幽默理解,同时将幽默生成视为新兴的下游前沿方向,并围绕多模态对齐、证据 grounded 推理与可控生成,对基准设计、评估协议与建模范式进行系统综述。This survey focuses on visual humor understanding in single-image and multi-panel artifacts, while treating humor generation as an emerging downstream frontier, and synthesizes benchmark design, evaluation protocols, and modeling paradigms based on multimodal alignment, evidence-grounded reasoning, and controlled generation.

Generative Adversarial Networks in Computer Vision: A Survey and Taxonomy
计算机视觉中的生成对抗网络:综述与分类
arXiv:1906.01529 多模态 综述 OA · 绿色 被引 277 · OpenAlex

深入回顾文献中 GAN 相关研究,并从两个视角阐述针对三大挑战所提出的架构变体与损失变体。An in-depth review of GAN-related research in the literature is provided, and an account of the architecture-variant and loss-variants, which have been proposed to handle these three challenges from two perspectives are provided.

Weights or Skills? A Survey of Robot-Learning Techniques: from Action-Predicting Weights to Robots that Write their Own Skills
Weights 还是 Skills?机器人学习方法综述:从预测动作的权重到自主编写技能的机器人
arXiv:2608.01851 多模态 综述 被引 0 · S2

该综述围绕"权重"与"技能"这一轴线组织领域,梳理了互补的"技能"一极——从无监督强化学习的技能发现,到大语言模型的技能库——并指出"skill"一词至少存在五种不同含义。This survey organises the field around that axis of weights versus skills, and maps the complementary"skills"pole, from unsupervised reinforcement-learning skill discovery to large-language-model skill libraries, and shows that the word "skill" is used in at least five distinct senses.

Domain Adaptation for Visual Applications: A Comprehensive Survey
视觉应用中的域适应:综合综述
arXiv:1702.05374 多模态 综述 OA · 绿色 被引 554 · S2

综述领域自适应与迁移学习,重点关注视觉应用及超越图像分类的方法,如目标检测、图像分割、视频分析或视觉属性学习。An overview of domain adaptation and transfer learning with a specific view on visual applications and the methods that go beyond image categorization, such as object detection or image segmentation, video analyses or learning visual attributes are overviewed.