Florence: A New Foundation Model for Computer Vision

  • 类型:arxiv
  • 标识:2111.11432
  • 链接:https://arxiv.org/abs/2111.11432
  • 主题:database
  • 主分类:multimodal
  • 形态:method
  • 被引:1152
  • 被引来源:Semantic Scholar
  • S2被引:1152
  • OpenAlex被引:342
  • 影响力被引:49
  • TLDR:This work introduces a new computer vision foundation model, Florence, to expand the representations from coarse (scene) to fine, from static (images) to dynamic (videos), and from RGB to multiple modalities (caption, depth), by incorporating universal visual-language representations from Web-scale image-text data.
  • OpenAlex ID:W3215626407
  • OpenAlex DOI:10.48550/arxiv.2111.11432
  • DOI:10.48550/arxiv.2111.11432
  • DOI来源:OpenAlex
  • 开放获取:green
  • 开放获取链接:https://arxiv.org/pdf/2111.11432
  • OpenAlex更新:2026-08-19
  • 待LLM分类:否
  • 标题中文:Florence:面向计算机视觉的新基础模型
  • TLDR中文:本文提出新的计算机视觉基础模型 Florence,通过融入来自 Web 规模图文数据的通用视觉-语言表示,将表征范围从粗粒度(场景)扩展到细粒度、从静态(图像)扩展到动态(视频)、从 RGB 扩展到多种模态(描述、深度等)。
  • 来源文件
  • [OpenAlex discover]
  • [S2 enrich]