VideoRAE: Taming Video Foundation Models for Generative Modeling via Representation Autoencoders

  • 类型:arxiv
  • 标识:2607.14088
  • 链接:https://arxiv.org/abs/2607.14088
  • 主分类:multimodal
  • 形态:method
  • 被引:0
  • 被引来源:Semantic Scholar
  • S2被引:0
  • 影响力被引:0
  • TLDR:VideoRAE is a representation autoencoder that leverages multi-scale hierarchical features from a frozen video foundation encoder and compresses them with a lightweight 1D self-attention projector to validate frozen VFM representations as versatile and generation-friendly video latents.
  • 待LLM分类:否
  • 标题中文:VideoRAE:通过表征自编码器驯服视频基础模型用于生成建模
  • TLDR中文:VideoRAE是一种表征自编码器,利用冻结视频基础编码器的多尺度分层特征,并通过轻量级1D自注意力投影器进行压缩,验证了冻结VFM表征可作为通用且利于生成的视频潜变量。
  • 来源文件
  • /inbox/tom/_candidates/2026-07-20-agent-rag-longcontext-candidates.json
  • [S2 enrich]