Flamingo: a Visual Language Model for Few-Shot Learning
- 类型:arxiv
- 标识:2204.14198
- 链接:https://arxiv.org/abs/2204.14198
- 主题:engineering
- 主分类:multimodal
- 形态:method
- 被引:6453
- 被引来源:Semantic Scholar
- S2被引:6453
- OpenAlex被引:1250
- 影响力被引:420
- TLDR:This work introduces Flamingo, a family of Visual Language Models (VLM) with this ability to bridge powerful pretrained vision-only and language-only models, handle sequences of arbitrarily interleaved visual and textual data, and seamlessly ingest images or videos as inputs.
- OpenAlex ID:W4225323055
- OpenAlex DOI:10.48550/arxiv.2204.14198
- DOI:10.48550/arxiv.2204.14198
- DOI来源:OpenAlex
- 开放获取:green
- 开放获取链接:https://arxiv.org/pdf/2204.14198
- OpenAlex更新:2026-08-26
- 待LLM分类:否
- 标题中文:Flamingo: a Visual Language Model for Few-Shot Learning
- 成熟度:production
- 场景:视觉语言模型、少样本学习
- TLDR中文:提出 Flamingo,一个 VLM 系列,能够桥接强大的纯视觉与纯语言预训练模型,处理任意交错排列的图文序列,并无缝接收图像或视频作为输入。
- 来源文件:
- [OpenAlex discover]
- [OpenAlex backfill]
- [S2 enrich]