On the Diffusibility of High-Dimensional Latents

  • 类型:arxiv
  • 标识:2609.28473
  • 链接:https://arxiv.org/abs/2609.28473
  • 主分类:engineering
  • 形态:method
  • TLDR:Representation Autoencoders (RAEs) enable diffusion models to operate in the feature spaces of pretrained visual encoders. However, many off-the-shelf encoders are not optimized for faithful reconstruction, discarding fine-grained visual details. As expected, finetuning these encoders for image reconstruction recovers such details. However, perhaps counterintuitively, this procedure reduces the effective dimensionality of the resulting representation, and the altered geometry has downstream effects on generation. Specifically, we show that using the standard velocity prediction in flow matchin
  • 待LLM分类:否
  • 来源文件
  • /inbox/tom/_candidates/2026-09-24-agent-rag-longcontext-candidates.json