Hugging Face Daily Papers · · 4 min read

MirrorWorld: Taming Video Diffusion Models for Mirror Reflection Generation

Mirrored from Hugging Face Daily Papers for archival readability. Support the source by reading on the original site.

<a href=\"https://cdn-uploads.huggingface.co/production/uploads/633aae379f5846dc43fe3af3/VYV9BqPrfvKebSwsMprvB.gif\" rel=\"nofollow\"><img src=\"https://cdn-uploads.huggingface.co/production/uploads/633aae379f5846dc43fe3af3/VYV9BqPrfvKebSwsMprvB.gif\" alt=\"mirrorworld-hero-mosaic\"></a></p>\n","updatedAt":"2026-08-11T17:40:26.936Z","author":{"_id":"633aae379f5846dc43fe3af3","avatarUrl":"/avatars/e4f71cc9a9357e0420653c05453ef8f7.svg","fullname":"YoujunZhao","name":"jhhjhhj","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.3718489408493042},"editors":["jhhjhhj"],"editorAvatarUrls":["/avatars/e4f71cc9a9357e0420653c05453ef8f7.svg"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.07463","authors":[{"_id":"6a7b37fbb7183653340c12ee","user":{"_id":"633aae379f5846dc43fe3af3","avatarUrl":"/avatars/e4f71cc9a9357e0420653c05453ef8f7.svg","isPro":false,"fullname":"YoujunZhao","user":"jhhjhhj","type":"user","name":"jhhjhhj"},"name":"Youjun Zhao","status":"claimed_verified","statusLastChangedAt":"2026-08-11T16:45:04.724Z","hidden":false},{"_id":"6a7b37fbb7183653340c12ef","name":"Alex Warren","hidden":false},{"_id":"6a7b37fbb7183653340c12f0","name":"Gary K. L. Tam","hidden":false},{"_id":"6a7b37fbb7183653340c12f1","name":"Rynson W. H. Lau","hidden":false}],"publishedAt":"2026-08-07T00:00:00.000Z","submittedOnDailyAt":"2026-08-11T00:00:00.000Z","title":"MirrorWorld: Taming Video Diffusion Models for Mirror Reflection Generation","submittedOnDailyBy":{"_id":"633aae379f5846dc43fe3af3","avatarUrl":"/avatars/e4f71cc9a9357e0420653c05453ef8f7.svg","isPro":false,"fullname":"YoujunZhao","user":"jhhjhhj","type":"user","name":"jhhjhhj"},"summary":"Recent advances in video diffusion models (VDMs) have enabled high-fidelity video synthesis. However, generating mirror reflections remains challenging because the content within a mirror must remain consistent with the surrounding scene. Existing VDMs are not specifically designed to model scene-to-mirror relationships, which can lead to reflections with incorrect content or inconsistent spatial arrangements. We observe that mirror reflection generation involves two complementary challenges: determining what scene content should be reflected and how the reflected content should be spatially arranged within the mirror region. Motivated by this observation, we propose MirrorWorld, a reflection-aware video inpainting framework that models scene-to-mirror relationships during generation. Specifically, we introduce Semantic Relation Distillation (SRD), which transfers relational information from a frozen visual foundation model to encourage semantic associations between visible scene content and mirror regions. We further propose Geometric Transformation Alignment (GTA), which learns a transformation that guides the spatial arrangement of reflected content. The two components play complementary roles, with SRD modeling what should be reflected and GTA modeling how it should be arranged. To facilitate research on this problem, we construct a benchmark for video mirror reflection generation by repurposing four existing video mirror datasets into a unified reflection reconstruction task. Experimental results show that MirrorWorld achieves improved reflection reconstruction quality over representative image-based reflection generation methods and strong video inpainting baselines.","upvotes":1,"discussionId":"6a7b37fbb7183653340c12f2","projectPage":"https://youjunzhao.github.io/MirrorWorld/","githubRepo":"https://github.com/YoujunZhao/MirrorWorld","githubRepoAddedBy":"user","ai_summary":"MirrorWorld improves video mirror reflection synthesis by separately modeling semantic content associations and geometric spatial arrangements through relation distillation and transformation alignment.","ai_keywords":["video diffusion models","mirror reflection generation","scene-to-mirror relationships","Semantic Relation Distillation","Geometric Transformation Alignment","video inpainting"],"ai_summary_model":"thinkingmachines/Inkling-Small","githubStars":6,"organization":{"_id":"65ffb2a4d2a378a163d5c459","name":"CityUniversityofHongKong","fullname":"City University of Hong Kong","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/65ffb1fca5242cffd5f83d60/fhD34yEelPToQ5RsLjeT0.png"}},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"633aae379f5846dc43fe3af3","avatarUrl":"/avatars/e4f71cc9a9357e0420653c05453ef8f7.svg","isPro":false,"fullname":"YoujunZhao","user":"jhhjhhj","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"organization":{"_id":"65ffb2a4d2a378a163d5c459","name":"CityUniversityofHongKong","fullname":"City University of Hong Kong","avatar":"https://cdn-avatars.huggingface.co/v1/production/uploads/65ffb1fca5242cffd5f83d60/fhD34yEelPToQ5RsLjeT0.png"},"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.07463.md","query":{}}">
Papers
arxiv:2608.07463

MirrorWorld: Taming Video Diffusion Models for Mirror Reflection Generation

Published on Aug 7
· Submitted by
YoujunZhao
on Aug 11
Authors:

Abstract

MirrorWorld improves video mirror reflection synthesis by separately modeling semantic content associations and geometric spatial arrangements through relation distillation and transformation alignment.

Recent advances in video diffusion models (VDMs) have enabled high-fidelity video synthesis. However, generating mirror reflections remains challenging because the content within a mirror must remain consistent with the surrounding scene. Existing VDMs are not specifically designed to model scene-to-mirror relationships, which can lead to reflections with incorrect content or inconsistent spatial arrangements. We observe that mirror reflection generation involves two complementary challenges: determining what scene content should be reflected and how the reflected content should be spatially arranged within the mirror region. Motivated by this observation, we propose MirrorWorld, a reflection-aware video inpainting framework that models scene-to-mirror relationships during generation. Specifically, we introduce Semantic Relation Distillation (SRD), which transfers relational information from a frozen visual foundation model to encourage semantic associations between visible scene content and mirror regions. We further propose Geometric Transformation Alignment (GTA), which learns a transformation that guides the spatial arrangement of reflected content. The two components play complementary roles, with SRD modeling what should be reflected and GTA modeling how it should be arranged. To facilitate research on this problem, we construct a benchmark for video mirror reflection generation by repurposing four existing video mirror datasets into a unified reflection reconstruction task. Experimental results show that MirrorWorld achieves improved reflection reconstruction quality over representative image-based reflection generation methods and strong video inpainting baselines.

Community

Paper author Paper submitter about 1 hour ago

mirrorworld-hero-mosaic

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images

· Sign up or log in to comment

Get this paper in your agent:

hf papers read 2608.07463
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper

No model linking this paper

Cite arxiv.org/abs/2608.07463 in a model README.md to link it from this page.

Datasets citing this paper

No dataset linking this paper

Cite arxiv.org/abs/2608.07463 in a dataset README.md to link it from this page.

Spaces citing this paper

No Space linking this paper

Cite arxiv.org/abs/2608.07463 in a Space README.md to link it from this page.

Collections including this paper

No Collection including this paper

Add this paper to a collection to link it from this page.

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hugging Face Daily Papers