UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos

  • 类型:arxiv
  • 标识:2608.11752
  • 链接:https://arxiv.org/abs/2608.11752
  • 主分类:multimodal
  • 形态:method
  • 被引:0
  • 被引来源:Semantic Scholar
  • S2被引:0
  • 影响力被引:0
  • TLDR:This work presents UniSwap, the first framework for streaming joint audio-visual identity replacement in talking videos, and introduces a swap-and-reconstruct pipeline that removes visual and vocal identity from real clips and uses the original clips as reconstruction targets.
  • 待LLM分类:否
  • 标题中文:UniSwap:面向说话视频的流式音视频身份替换
  • TLDR中文:本文提出 UniSwap,这是首个用于说话视频中流式联合音视频身份替换的框架,并引入 swap-and-reconstruct 流程,从真实片段中移除视觉和声音身份,同时使用原始片段作为重建目标。
  • 来源文件
  • /inbox/tom/_candidates/2026-08-14-agent-rag-longcontext-candidates.json
  • [S2 enrich]