PreviewDiff: Multimodal Critic-Guided Search over Diffusion Latents

  • 类型:arxiv
  • 标识:2609.36199
  • 链接:https://arxiv.org/abs/2609.36199
  • 主分类:multimodal
  • 形态:position
  • TLDR:Diffusion models can produce striking images and videos, but they still struggle with the compositional details that make a generation faithful to a prompt, such as object counts, attribute binding, spatial relations, and temporally grounded actions. A common way to improve prompt satisfaction is to spend more compute at test time through Best-of-N sampling, but final-sample selection is fixed. Best-of-N can only choose among completed outputs and cannot repair a promising trajectory before it fails. We introduce PreviewDiff, a training-free test-time search method that turns diffusion samplin
  • 待LLM分类:否
  • 来源文件:
  • /inbox/tom/_candidates/2026-10-01-agent-rag-longcontext-candidates.json