UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos
- 类型:arxiv
- 标识:2608.11752
- 链接:https://arxiv.org/abs/2608.11752
- 主分类:multimodal
- 形态:method
- 被引:0
- 被引来源:Semantic Scholar + OpenAlex
- S2被引:0
- OpenAlex被引:0
- 影响力被引:0
- TLDR:This work presents UniSwap, the first framework for streaming joint audio-visual identity replacement in talking videos, and introduces a swap-and-reconstruct pipeline that removes visual and vocal identity from real clips and uses the original clips as reconstruction targets.
- OpenAlex ID:W7202349423
- OpenAlex DOI:10.48550/arxiv.2608.11752
- DOI:10.48550/arxiv.2608.11752
- DOI来源:OpenAlex
- 开放获取:green
- 开放获取链接:https://arxiv.org/pdf/2608.11752
- OpenAlex更新:2026-08-27
- 待LLM分类:否
- 标题中文:UniSwap:面向说话视频的流式音视频身份替换
- TLDR中文:本文提出 UniSwap,这是首个用于说话视频中流式联合音视频身份替换的框架,并引入 swap-and-reconstruct 流程,从真实片段中移除视觉和声音身份,同时使用原始片段作为重建目标。
- 来源文件:
- /inbox/tom/_candidates/2026-08-14-agent-rag-longcontext-candidates.json
- [S2 enrich]
- [OpenAlex backfill]