Video Generation Models: A Survey of Post-Training and Alignment

  • 类型:arxiv
  • 标识:2610.00812
  • 链接:https://arxiv.org/abs/2610.00812
  • 主分类:multimodal
  • 形态:survey
  • TLDR:Video generation has rapidly progressed from short, low-quality clips to high-resolution, long-duration sequences with complex spatiotemporal dynamics. Despite strong generative priors learned through large-scale pretraining, pretrained video models often fail to reliably follow human intent, maintain temporal coherence, or satisfy physical and safety constraints. Compared with image and text generation, alignment in video generation presents unique challenges, including error accumulation over time, motion-appearance coupling, multi-objective trade-offs, and limited supervision for temporal p
  • 副分类:engineering
  • 待LLM分类:否
  • 标题中文:视频生成模型:后训练与对齐综述
  • TLDR中文:视频生成已从短小的低质量片段迅速发展为具有复杂时空动态的高分辨率长时长序列。尽管通过大规模预训练学到了强大的生成先验,但预训练视频模型往往无法可靠地遵循人类意图、保持时间一致性,或满足物理与安全约束。与图像和文本生成相比,视频生成中的对齐面临独特挑战,包括随时间的误差累积、运动-外观耦合、多目标权衡,以及对时间合理性的监督有限。
  • 来源文件:
  • /inbox/tom/_candidates/2026-10-03-agent-rag-longcontext-candidates.json