Deep Visual-Semantic Alignments for Generating Image Descriptions

  • 类型:arxiv
  • 标识:1412.2306
  • 链接:https://arxiv.org/abs/1412.2306
  • 主题:database
  • 主分类:risk
  • 形态:method
  • 被引:6111
  • 被引来源:Semantic Scholar
  • S2被引:6111
  • OpenAlex被引:148
  • 影响力被引:513
  • TLDR:A model that generates natural language descriptions of images and their regions based on a novel combination of Convolutional Neural Networks over image regions, bidirectional Recurrent Neural Networks over sentences, and a structured objective that aligns the two modalities through a multimodal embedding is presented.
  • OpenAlex ID:W2951805548
  • OpenAlex DOI:10.48550/arxiv.1412.2306
  • DOI:10.48550/arxiv.1412.2306
  • DOI来源:OpenAlex
  • 开放获取:green
  • 开放获取链接:https://arxiv.org/pdf/1412.2306
  • OpenAlex更新:2026-08-03
  • 待LLM分类:否
  • 标题中文:用于生成图像描述的深度视觉-语义对齐
  • TLDR中文:提出一个模型,基于图像区域上的 CNN、句子上的双向 RNN 以及通过多模态嵌入对齐两种模态的结构化目标,生成图像及其区域的自然语言描述。
  • 来源文件
  • [OpenAlex discover]
  • [S2 enrich]