BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation
- 类型:arxiv
- 标识:2201.12086
- 链接:https://arxiv.org/abs/2201.12086
- 主题:multimodal
- 主分类:multimodal
- 形态:method
- 被引:7222
- 被引来源:Semantic Scholar
- S2被引:7222
- OpenAlex被引:870
- 影响力被引:749
- TLDR:BLIP effectively utilizes the noisy web data by bootstrapping the captions, where a captioner generates synthetic captions and a filter removes the noisy ones, and demonstrates strong generalization ability when directly transferred to video-language tasks in a zero-shot manner.
- OpenAlex ID:W4226182655
- OpenAlex DOI:10.48550/arxiv.2201.12086
- DOI:10.48550/arxiv.2201.12086
- DOI来源:OpenAlex
- 开放获取:green
- 开放获取链接:https://arxiv.org/pdf/2201.12086
- OpenAlex更新:2026-08-23
- 副分类:engineering
- 待LLM分类:否
- 标题中文:BLIP: 面向统一视觉-语言理解与生成的引导式语言-图像预训练
- TLDR中文:BLIP 通过引导式 caption 方式有效利用含噪网络数据,由 captioner 生成合成 caption,并由 filter 去除噪声样本;在以零样本方式直接迁移到视频-语言任务时,展现出强大的泛化能力。
- 来源文件:
- [OpenAlex discover]
- [S2 enrich]