We present UniSwap, a framework for streaming joint audio-visual identity replacement in talking videos. Unlike existing methods that optimize appearance and voice using separate models, UniSwap performs joint transfer within a single audio-visual diffusion transformer to ensure multi-modal consistency. To address training data scarcity, we use a swap-and-reconstruct pipeline that extracts identities from real clips and targets the original media for reconstruction. Our approach utilizes In-context Pretraining for joint replacement, Conditional Streaming Adaptation for block-causal KV-cached generation, and an Efficient Self-forcing DMD mechanism that limits sampling to 3 denoising steps per block. Combined with Feature-RoPE Decomposition for stable long-form inference, UniSwap achieves highly synchronized, streaming-efficient audio-visual identity replacement.</p>\n","updatedAt":"2026-08-14T05:15:32.143Z","author":{"_id":"64489bb5e21484883408a96d","avatarUrl":"/avatars/23f7aa4d733a7708fab4ee059ac2b323.svg","fullname":"Jinpeng YU","name":"Jacob-Yu","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":3,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.815604031085968},"editors":["Jacob-Yu"],"editorAvatarUrls":["/avatars/23f7aa4d733a7708fab4ee059ac2b323.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.11752","authors":[{"_id":"6a7ea07b42823931a1f1773b","name":"Yuxuan Zhang","hidden":false},{"_id":"6a7ea07b42823931a1f1773c","name":"Haozhong Xiong","hidden":false},{"_id":"6a7ea07b42823931a1f1773d","name":"Jiayi Song","hidden":false},{"_id":"6a7ea07b42823931a1f1773e","user":{"_id":"64489bb5e21484883408a96d","avatarUrl":"/avatars/23f7aa4d733a7708fab4ee059ac2b323.svg","isPro":false,"fullname":"Jinpeng YU","user":"Jacob-Yu","type":"user","name":"Jacob-Yu"},"name":"Jinpeng Yu","status":"claimed_verified","statusLastChangedAt":"2026-08-14T08:45:04.777Z","hidden":false},{"_id":"6a7ea07b42823931a1f1773f","name":"Yang Shi","hidden":false},{"_id":"6a7ea07b42823931a1f17740","name":"Jiaming Liu","hidden":false},{"_id":"6a7ea07b42823931a1f17741","name":"Ruihua Huang","hidden":false},{"_id":"6a7ea07b42823931a1f17742","name":"Liwei Wang","hidden":false}],"publishedAt":"2026-08-13T00:00:00.000Z","submittedOnDailyAt":"2026-08-14T00:00:00.000Z","title":"UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos","submittedOnDailyBy":{"_id":"64489bb5e21484883408a96d","avatarUrl":"/avatars/23f7aa4d733a7708fab4ee059ac2b323.svg","isPro":false,"fullname":"Jinpeng YU","user":"Jacob-Yu","type":"user","name":"Jacob-Yu"},"summary":"Talking-video character replacement requires coordinated transfer of appearance and voice while preserving the source motion, scene, linguistic content, and audio-video timing. Existing methods use separately optimized models for the two modalities, making audio-visual consistency difficult to enforce. We present UniSwap, the first framework for streaming joint audio-visual identity replacement in talking videos. Given a source video, a reference image, and a reference voice clip, UniSwap transfers the reference appearance and vocal timbre within a single audio-visual diffusion transformer while preserving the source content and dynamics. To address the scarcity of aligned cross-identity training pairs, we introduce a swap-and-reconstruct pipeline that removes visual and vocal identity from real clips and uses the original clips as reconstruction targets. Starting from a bidirectional backbone, we progressively adapt the model through In-context Pretraining for joint replacement, Conditional Streaming Adaptation for block-causal KV-cached generation, and Efficient Self-forcing DMD for mitigating exposure bias and reducing sampling from 30 to 3 denoising steps per block. Efficient Multi-LoRA Switching enables the three DMD roles to share a single frozen backbone. Feature-RoPE Decomposition keeps cached positions within the training range, supporting stable long-form inference. Experiments demonstrate strong audio-visual synchronization, competitive identity preservation, efficient streaming, and stable long-form generation.","upvotes":10,"discussionId":"6a7ea07b42823931a1f17743","projectPage":"https://uniswap-av.github.io/","githubRepo":"https://github.com/uniswap-av/UniSwap","githubRepoAddedBy":"user","ai_summary":"UniSwap enables synchronized appearance and voice replacement in talking videos through a unified streaming audio-visual diffusion transformer with specialized training and inference adaptations.","ai_keywords":["audio-visual diffusion transformer","swap-and-reconstruct pipeline","In-context Pretraining","Conditional Streaming Adaptation","block-causal KV-cached generation","Efficient Self-forcing DMD","Multi-LoRA Switching","Feature-RoPE Decomposition","denoising steps"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":4,"organization":{"_id":"6a6841e7107886ba1a151b03","name":"QwenBusinessUnit","fullname":"Qwen Business Unit","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/66f79b323fe089b75e9e0c04/MlefZsdry-JuhKzAwxjQl.webp"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"64489bb5e21484883408a96d","avatarUrl":"/avatars/23f7aa4d733a7708fab4ee059ac2b323.svg","isPro":false,"fullname":"Jinpeng YU","user":"Jacob-Yu","type":"user"},{"_id":"66d53350ad293ffc4b178d10","avatarUrl":"/avatars/d1cb7c800d2af4081de62ef95dc2c44c.svg","isPro":false,"fullname":"jackylova","user":"jackylova","type":"user"},{"_id":"688cceac29c69ec01eb28644","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/wyL1U86pB5GIi4x3JflOa.png","isPro":false,"fullname":"Chuyue Li","user":"woody-woody","type":"user"},{"_id":"67051ad602b48c3fce373fe7","avatarUrl":"/avatars/ebf10382f7c34f41b77f30eb1eea784e.svg","isPro":false,"fullname":"sjy","user":"sjy92","type":"user"},{"_id":"636b3f9ce3ad78bc68b67541","avatarUrl":"/avatars/2b7e745953ae39e01222e99fb63b279e.svg","isPro":false,"fullname":"yuxuan","user":"zzyx","type":"user"},{"_id":"68f9a5cf4cbb5261e6a3cf78","avatarUrl":"/avatars/a0e9bcadd2b69e35ccca6f7d75c2b0b0.svg","isPro":false,"fullname":"Ethan Miller","user":"YaKaYC","type":"user"},{"_id":"692d7843a918d2fbb2893b15","avatarUrl":"/avatars/d68856827c5f1cd0aecc28398fc84757.svg","isPro":false,"fullname":"yang","user":"Newsyshi","type":"user"},{"_id":"69be48e4df85f5f792fc2a7e","avatarUrl":"/avatars/06505fd14a44973c6305ec3f23a67ee4.svg","isPro":false,"fullname":"None","user":"Gonzalo320","type":"user"},{"_id":"65f30652b0d359b2ffa4a42c","avatarUrl":"/avatars/f6cb0705b25cc0ccc0b099811ef9a871.svg","isPro":false,"fullname":"zrx","user":"z-rx","type":"user"},{"_id":"620783f24e28382272337ba4","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/620783f24e28382272337ba4/zkUveQPNiDfYjgGhuFErj.jpeg","isPro":false,"fullname":"GuoLiangTang","user":"Tommy930","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"6a6841e7107886ba1a151b03","name":"QwenBusinessUnit","fullname":"Qwen Business Unit","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/66f79b323fe089b75e9e0c04/MlefZsdry-JuhKzAwxjQl.webp"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.11752.md","query":{}}">
UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos
Abstract
UniSwap enables synchronized appearance and voice replacement in talking videos through a unified streaming audio-visual diffusion transformer with specialized training and inference adaptations.
Talking-video character replacement requires coordinated transfer of appearance and voice while preserving the source motion, scene, linguistic content, and audio-video timing. Existing methods use separately optimized models for the two modalities, making audio-visual consistency difficult to enforce. We present UniSwap, the first framework for streaming joint audio-visual identity replacement in talking videos. Given a source video, a reference image, and a reference voice clip, UniSwap transfers the reference appearance and vocal timbre within a single audio-visual diffusion transformer while preserving the source content and dynamics. To address the scarcity of aligned cross-identity training pairs, we introduce a swap-and-reconstruct pipeline that removes visual and vocal identity from real clips and uses the original clips as reconstruction targets. Starting from a bidirectional backbone, we progressively adapt the model through In-context Pretraining for joint replacement, Conditional Streaming Adaptation for block-causal KV-cached generation, and Efficient Self-forcing DMD for mitigating exposure bias and reducing sampling from 30 to 3 denoising steps per block. Efficient Multi-LoRA Switching enables the three DMD roles to share a single frozen backbone. Feature-RoPE Decomposition keeps cached positions within the training range, supporting stable long-form inference. Experiments demonstrate strong audio-visual synchronization, competitive identity preservation, efficient streaming, and stable long-form generation.
Community
We present UniSwap, a framework for streaming joint audio-visual identity replacement in talking videos. Unlike existing methods that optimize appearance and voice using separate models, UniSwap performs joint transfer within a single audio-visual diffusion transformer to ensure multi-modal consistency. To address training data scarcity, we use a swap-and-reconstruct pipeline that extracts identities from real clips and targets the original media for reconstruction. Our approach utilizes In-context Pretraining for joint replacement, Conditional Streaming Adaptation for block-causal KV-cached generation, and an Efficient Self-forcing DMD mechanism that limits sampling to 3 denoising steps per block. Combined with Feature-RoPE Decomposition for stable long-form inference, UniSwap achieves highly synchronized, streaming-efficient audio-visual identity replacement.
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2608.11752 in a model README.md to link it from this page.
Cite arxiv.org/abs/2608.11752 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2608.11752 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.