BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation

  • 类型:arxiv
  • 标识:2201.12086
  • 链接:https://arxiv.org/abs/2201.12086
  • 主题:multimodal
  • 主分类:multimodal
  • 形态:method
  • 被引:7222
  • 被引来源:Semantic Scholar
  • S2被引:7222
  • OpenAlex被引:870
  • 影响力被引:749
  • TLDR:BLIP effectively utilizes the noisy web data by bootstrapping the captions, where a captioner generates synthetic captions and a filter removes the noisy ones, and demonstrates strong generalization ability when directly transferred to video-language tasks in a zero-shot manner.
  • OpenAlex ID:W4226182655
  • OpenAlex DOI:10.48550/arxiv.2201.12086
  • DOI:10.48550/arxiv.2201.12086
  • DOI来源:OpenAlex
  • 开放获取:green
  • 开放获取链接:https://arxiv.org/pdf/2201.12086
  • OpenAlex更新:2026-08-23
  • 副分类:engineering
  • 待LLM分类:否
  • 标题中文:BLIP: 面向统一视觉-语言理解与生成的引导式语言-图像预训练
  • TLDR中文:BLIP 通过引导式 caption 方式有效利用含噪网络数据,由 captioner 生成合成 caption,并由 filter 去除噪声样本;在以零样本方式直接迁移到视频-语言任务时,展现出强大的泛化能力。
  • 来源文件
  • [OpenAlex discover]
  • [S2 enrich]