Flamingo: a Visual Language Model for Few-Shot Learning

  • 类型:arxiv
  • 标识:2204.14198
  • 链接:https://arxiv.org/abs/2204.14198
  • 主题:engineering
  • 主分类:multimodal
  • 形态:method
  • 被引:6453
  • 被引来源:Semantic Scholar
  • S2被引:6453
  • OpenAlex被引:1250
  • 影响力被引:420
  • TLDR:This work introduces Flamingo, a family of Visual Language Models (VLM) with this ability to bridge powerful pretrained vision-only and language-only models, handle sequences of arbitrarily interleaved visual and textual data, and seamlessly ingest images or videos as inputs.
  • OpenAlex ID:W4225323055
  • OpenAlex DOI:10.48550/arxiv.2204.14198
  • DOI:10.48550/arxiv.2204.14198
  • DOI来源:OpenAlex
  • 开放获取:green
  • 开放获取链接:https://arxiv.org/pdf/2204.14198
  • OpenAlex更新:2026-08-26
  • 待LLM分类:否
  • 标题中文:Flamingo: a Visual Language Model for Few-Shot Learning
  • 成熟度:production
  • 场景:视觉语言模型、少样本学习
  • TLDR中文:提出 Flamingo,一个 VLM 系列,能够桥接强大的纯视觉与纯语言预训练模型,处理任意交错排列的图文序列,并无缝接收图像或视频作为输入。
  • 来源文件
  • [OpenAlex discover]
  • [OpenAlex backfill]
  • [S2 enrich]