Deep Captioning with Multimodal Recurrent Neural Networks (m-RNN)

  • 类型:arxiv
  • 标识:1412.6632
  • 链接:https://arxiv.org/abs/1412.6632
  • 主题:multimodal
  • 主分类:multimodal
  • 形态:method
  • 被引:1283
  • 被引来源:Semantic Scholar
  • S2被引:1283
  • OpenAlex被引:652
  • 影响力被引:114
  • TLDR:The m-RNN model directly models the probability distribution of generating a word given previous words and an image, and achieves significant performance improvement over the state-of-the-art methods which directly optimize the ranking objective function for retrieval.
  • OpenAlex ID:W1811254738
  • OpenAlex DOI:10.48550/arxiv.1412.6632
  • DOI:10.48550/arxiv.1412.6632
  • DOI来源:OpenAlex
  • 开放获取:green
  • 开放获取链接:https://arxiv.org/pdf/1412.6632
  • OpenAlex更新:2026-08-23
  • 待LLM分类:否
  • 标题中文:Deep Captioning with Multimodal Recurrent Neural Networks (m-RNN)
  • TLDR中文:m-RNN 模型直接对给定先前词语和图像条件下生成下一个词的概率分布建模,相较于直接优化排序目标函数进行检索的 SOTA 方法,取得了显著的性能提升。
  • 来源文件
  • [OpenAlex discover]
  • [S2 enrich]