InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning
- 类型:arxiv
- 标识:2305.06500
- 链接:https://arxiv.org/abs/2305.06500
- 主题:risk
- 主分类:multimodal
- 形态:method
- 被引:3878
- 被引来源:Semantic Scholar
- S2被引:3878
- OpenAlex被引:409
- 影响力被引:555
- TLDR:This paper conducts a systematic and comprehensive study on vision-language instruction tuning based on the pretrained BLIP-2 models, and introduces an instruction-aware Query Transformer, which extracts informative features tailored to the given instruction.
- OpenAlex ID:W4376312115
- OpenAlex DOI:10.48550/arxiv.2305.06500
- DOI:10.48550/arxiv.2305.06500
- DOI来源:OpenAlex
- 开放获取:green
- 开放获取链接:https://arxiv.org/pdf/2305.06500
- OpenAlex更新:2026-08-20
- 待LLM分类:否
- 标题中文:InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning
- TLDR中文:本文基于预训练 BLIP-2 模型,对视觉-语言指令微调展开系统全面研究,并提出指令感知的 Query Transformer,用于提取针对给定指令的信息丰富特征。
- 来源文件:
- [OpenAlex discover]
- [S2 enrich]