EffectLearner enables robust real-world video object removal by jointly removing target objects and their associated effects, including shadows, reflections, lighting changes, and other complex physical interactions.</p>\n","updatedAt":"2026-08-07T06:07:48.535Z","author":{"_id":"66a356f2f7352f4ffbb4e74a","avatarUrl":"/avatars/4391990768ffc079cf9146973614ba84.svg","fullname":"He","name":"Phoebux","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":1,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.9307298064231873},"editors":["Phoebux"],"editorAvatarUrls":["/avatars/4391990768ffc079cf9146973614ba84.svg"],"reactions":[],"isReport":false}},{"id":"6a7576574540b42216763b7a","author":{"_id":"66a356f2f7352f4ffbb4e74a","avatarUrl":"/avatars/4391990768ffc079cf9146973614ba84.svg","fullname":"He","name":"Phoebux","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":1,"isUserFollowing":false},"createdAt":"2026-08-07T06:08:23.000Z","type":"comment","data":{"edited":false,"hidden":false,"latest":{"raw":"We also introduce EffectWorld, a large-scale paired video dataset designed for effect-aware object removal, covering compositional effects, weak object-effect correlations, long-tail physical phenomena, and dynamic challenges: https://huggingface.co/datasets/jenniferwuuu/EffectWorld\n","html":"<p>We also introduce EffectWorld, a large-scale paired video dataset designed for effect-aware object removal, covering compositional effects, weak object-effect correlations, long-tail physical phenomena, and dynamic challenges: <a href=\"https://huggingface.co/datasets/jenniferwuuu/EffectWorld\">https://huggingface.co/datasets/jenniferwuuu/EffectWorld</a></p>\n","updatedAt":"2026-08-07T06:08:23.941Z","author":{"_id":"66a356f2f7352f4ffbb4e74a","avatarUrl":"/avatars/4391990768ffc079cf9146973614ba84.svg","fullname":"He","name":"Phoebux","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"followerCount":1,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.7902067303657532},"editors":["Phoebux"],"editorAvatarUrls":["/avatars/4391990768ffc079cf9146973614ba84.svg"],"reactions":[],"isReport":false}},{"id":"6a7592f490ade50401c9377f","author":{"_id":"677d638ad6ad451793e30c7f","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/677d638ad6ad451793e30c7f/IxUxdPsogK_M02asRw1kn.png","fullname":"jenniferwuuu","name":"jenniferwuuu","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false},"createdAt":"2026-08-07T08:10:28.000Z","type":"comment","data":{"edited":false,"hidden":false,"latest":{"raw":"\nhttps://cdn-uploads.huggingface.co/production/uploads/677d638ad6ad451793e30c7f/3TpQtMblQBCIvemqZtmzJ.mp4\n","html":"<p><video src=\"https://cdn-uploads.huggingface.co/production/uploads/677d638ad6ad451793e30c7f/3TpQtMblQBCIvemqZtmzJ.mp4\" controls=\"\" class=\"max-w-full!\"></video></p>\n","updatedAt":"2026-08-07T08:10:28.700Z","author":{"_id":"677d638ad6ad451793e30c7f","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/677d638ad6ad451793e30c7f/IxUxdPsogK_M02asRw1kn.png","fullname":"jenniferwuuu","name":"jenniferwuuu","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false}},"numEdits":0,"identifiedLanguage":{"language":"en","probability":0.5727272033691406},"editors":["jenniferwuuu"],"editorAvatarUrls":["https://cdn-avatars.huggingface.co/v1/production/uploads/677d638ad6ad451793e30c7f/IxUxdPsogK_M02asRw1kn.png"],"reactions":[],"isReport":false}}],"primaryEmailConfirmed":false,"paper":{"id":"2608.05565","authors":[{"_id":"6a75722ee1228e04b3238303","name":"Feier Wu","hidden":false},{"_id":"6a75722ee1228e04b3238304","name":"Wanke Xia","hidden":false},{"_id":"6a75722ee1228e04b3238305","name":"Xu He","hidden":false},{"_id":"6a75722ee1228e04b3238306","name":"Zilang Zhou","hidden":false},{"_id":"6a75722ee1228e04b3238307","name":"Si Chen","hidden":false},{"_id":"6a75722ee1228e04b3238308","name":"Dongxia Liu","hidden":false},{"_id":"6a75722ee1228e04b3238309","name":"Liyang Chen","hidden":false},{"_id":"6a75722ee1228e04b323830a","name":"Qimeng Wu","hidden":false},{"_id":"6a75722ee1228e04b323830b","name":"Zhengbo Zhang","hidden":false},{"_id":"6a75722ee1228e04b323830c","name":"Wenming Yang","hidden":false},{"_id":"6a75722ee1228e04b323830d","name":"Zhiyong Wu","hidden":false}],"publishedAt":"2026-08-06T00:00:00.000Z","submittedOnDailyAt":"2026-08-07T00:00:00.000Z","title":"EffectLearner: World-Aware Object-Effect Reasoning for Real-World Video Object Removal","submittedOnDailyBy":{"_id":"66a356f2f7352f4ffbb4e74a","avatarUrl":"/avatars/4391990768ffc079cf9146973614ba84.svg","isPro":false,"fullname":"He","user":"Phoebux","type":"user","name":"Phoebux"},"summary":"Video object removal must eliminate not only the target object but also its induced effects while maintaining high-fidelity and spatiotemporally coherent restoration. Existing methods mainly learn object-effect correspondences implicitly from predefined effect categories and fixed data distributions, limiting their generalization to complex real-world scenes involving compositional effects, spatially detached or weakly correlated effects, long-tail physical phenomena, and dynamically evolving interactions. We propose EffectLearner, a semantic-reasoning-enhanced framework that combines a VLM-based Object-Effect Reasoner with a DiT-based Video Eraser. Guided by a structured effect-analysis prompt, the Reasoner performs cross-modal reasoning over a target-highlighted video and extracts compact effect-aware context, which guides the Video Eraser toward comprehensive object-effect removal. Motion-aware mask guidance and motion-consistency supervision further improve removal coverage and spatiotemporal stability under object motion and evolving scene dynamics. To fully exploit the framework in challenging real-world scenarios, we further construct EffectWorld, a paired video dataset specifically designed for complex object-induced effects, and introduce a progressive training curriculum that combines common supervision with complex-effect data. On the standard ROSE-Bench, EffectLearner outperforms existing baselines on most metrics and achieves clear advantages on both EffectWorld-Eval and the challenging EffectWorld-Wild, demonstrating its ability to deliver high-quality video object removal in complex real-world scenes.","upvotes":12,"discussionId":"6a75722ee1228e04b323830e","projectPage":"https://morleyolsen.github.io/EffectLearner/"},"canReadDatabase":false,"canManagePapers":false,"canSubmit":false,"hasHfLevelAccess":false,"upvoted":false,"upvoters":[{"_id":"693d79b8e187be5c17b8c276","avatarUrl":"/avatars/e3b8c166f2e5a4ca61f6654276043062.svg","isPro":false,"fullname":"Wanke Xia","user":"xwk25","type":"user"},{"_id":"66a356f2f7352f4ffbb4e74a","avatarUrl":"/avatars/4391990768ffc079cf9146973614ba84.svg","isPro":false,"fullname":"He","user":"Phoebux","type":"user"},{"_id":"677d638ad6ad451793e30c7f","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/677d638ad6ad451793e30c7f/IxUxdPsogK_M02asRw1kn.png","isPro":false,"fullname":"jenniferwuuu","user":"jenniferwuuu","type":"user"},{"_id":"67fab241a0b4ccba867ef93c","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/no-auth/Dm7LsQd9ZPzFh71i_o0Ke.png","isPro":false,"fullname":"yanxt","user":"Yanxt-re","type":"user"},{"_id":"671906d1827bea3146f0b457","avatarUrl":"/avatars/4cf0b766143768e58f3cc41673f77cb1.svg","isPro":false,"fullname":"Bo","user":"Zhengbo-Zhang","type":"user"},{"_id":"64944333b14db30b0a96fafc","avatarUrl":"/avatars/c00425fbcd2a9509c2ed6e65fad5f6f1.svg","isPro":false,"fullname":"Xinyu Chen","user":"CocaCat","type":"user"},{"_id":"6715b493d54796e4b99d90e8","avatarUrl":"https://cdn-avatars.huggingface.co/v1/production/uploads/6715b493d54796e4b99d90e8/X-VqGsRrjPlfb94GRn13x.jpeg","isPro":false,"fullname":"黄炜锴","user":"tsrigo","type":"user"},{"_id":"656c3503e0ff1cebe95e5730","avatarUrl":"/avatars/e9e6ecde1ca79833bfe23df51ee0dc68.svg","isPro":false,"fullname":"Tsuki","user":"Tsukihjy","type":"user"},{"_id":"694a76be5ba6d97e99828360","avatarUrl":"/avatars/78d8f772a29572917ff49f36390b0cca.svg","isPro":false,"fullname":"Shiqing Meng","user":"qqz123","type":"user"},{"_id":"69019c7116bd45305a496baf","avatarUrl":"/avatars/cbb961532ca904befd14f385fdd7d0af.svg","isPro":false,"fullname":"Wanke Xia","user":"MorleyOlsen","type":"user"},{"_id":"673d455fbcc5f8535d73e46e","avatarUrl":"/avatars/ee1832fbac66b24e428ff3a4ff71f2f7.svg","isPro":false,"fullname":"king yuan","user":"lemon0703","type":"user"},{"_id":"6a75b0bd437a2a142e63175a","avatarUrl":"/avatars/4fb9b582283e4ef82b59aaa8403f6d9a.svg","isPro":false,"fullname":"Bing-Zhou Chen","user":"CBZ199671","type":"user"}],"acceptLanguages":["en"],"dailyPaperRank":0,"markdownContentUrl":"https://huggingface.co/buckets/huggingchat/papers-content/resolve/2608/2608.05565.md","query":{}}">
EffectLearner: World-Aware Object-Effect Reasoning for Real-World Video Object Removal
Published on Aug 6
· Submitted by He on Aug 7 Abstract
Video object removal must eliminate not only the target object but also its induced effects while maintaining high-fidelity and spatiotemporally coherent restoration. Existing methods mainly learn object-effect correspondences implicitly from predefined effect categories and fixed data distributions, limiting their generalization to complex real-world scenes involving compositional effects, spatially detached or weakly correlated effects, long-tail physical phenomena, and dynamically evolving interactions. We propose EffectLearner, a semantic-reasoning-enhanced framework that combines a VLM-based Object-Effect Reasoner with a DiT-based Video Eraser. Guided by a structured effect-analysis prompt, the Reasoner performs cross-modal reasoning over a target-highlighted video and extracts compact effect-aware context, which guides the Video Eraser toward comprehensive object-effect removal. Motion-aware mask guidance and motion-consistency supervision further improve removal coverage and spatiotemporal stability under object motion and evolving scene dynamics. To fully exploit the framework in challenging real-world scenarios, we further construct EffectWorld, a paired video dataset specifically designed for complex object-induced effects, and introduce a progressive training curriculum that combines common supervision with complex-effect data. On the standard ROSE-Bench, EffectLearner outperforms existing baselines on most metrics and achieves clear advantages on both EffectWorld-Eval and the challenging EffectWorld-Wild, demonstrating its ability to deliver high-quality video object removal in complex real-world scenes.
Community
EffectLearner enables robust real-world video object removal by jointly removing target objects and their associated effects, including shadows, reflections, lighting changes, and other complex physical interactions.
We also introduce EffectWorld, a large-scale paired video dataset designed for effect-aware object removal, covering compositional effects, weak object-effect correlations, long-tail physical phenomena, and dynamic challenges: https://huggingface.co/datasets/jenniferwuuu/EffectWorld
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
Cite arxiv.org/abs/2608.05565 in a model README.md to link it from this page.
Cite arxiv.org/abs/2608.05565 in a dataset README.md to link it from this page.
Cite arxiv.org/abs/2608.05565 in a Space README.md to link it from this page.
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.